GDS Technology — Built, Wired and Secured podcast banner
Watch on YouTube →
Episodes Built
Episode 88

Patch Windows: Coordinating BAS & OT Updates Without Breaking Operations

July 24, 2026
Key takeaways
  • BAS and OT patching must be planned around operational risk, not treated like standard endpoint updates.
  • A three-owner model creates clarity around the request, operational oversight, and rollback authority.
  • Five fast validation checks can catch tenant-impacting issues before a rollout spreads.
  • Rollback plans should start with small atomic reversions, not full image restores.
  • Tiered notifications and tabletop drills help reduce confusion during planned maintenance windows.

Show Notes

Why Patch Windows for BAS and OT Need a Different Playbook

This episode of Built, Wired & Secured opens with a scenario that feels familiar to anyone responsible for building operations: a routine Windows update hits front-end controllers at 2 a.m., and suddenly HVAC drops on tenant floors, card access turns unreliable, phones start ringing, and facilities is forced into a scramble. The bigger problem is not just the failed update. It is the lack of clarity around ownership, authority, and rollback.

The conversation centers on a simple but critical question: what breaks if this goes down? That framing shifts patching from a routine technical task to an operational risk decision. In BAS and OT environments, updates are not just about software hygiene. They affect comfort, access, scheduling, telemetry, and the tenant experience in real time.

How BAS and OT Patching Differs from Standard IT Updates

The guests explain that traditional IT endpoints are often more forgiving. Reboots can be tolerated, devices may be more disposable, and business disruption can be contained. BAS and OT systems are different because they sit inside persistent operational dependencies.

  • HVAC schedules may be interrupted by a reboot or service failure
  • Access control systems can become unreliable during update cycles
  • Elevator logic and other building functions may depend on stable controller behavior
  • Vendor interfaces and proprietary firmware create additional coordination challenges
  • Maintenance windows often conflict with real tenant and facility needs

Because of that, a one-size-fits-all patch window is described as a mistake. Some devices can be safely updated after hours, while others should first be staged in a lab or tested on a non-production floor. The takeaway is practical: test where it matters most rather than treating every device the same.

Why Teams Delay Necessary Updates

The episode also addresses why updates get deferred even when the security case is obvious. Fear of downtime is part of the answer, but the discussion points more directly to incentives. Immediate overtime, scheduling difficulty, and perceived disruption are visible costs. The risk of an unpatched vulnerability feels distant and abstract, so patches are postponed until the situation becomes urgent.

At the same time, the guests warn against chasing perfect validation across every possible device and integration before moving forward. They do not argue for less testing. They argue for prioritized testing. Critical tenant-impacting workflows should be validated first so teams can reduce risk without falling into paralysis.

A Lightweight Change Control Flow That Actually Works

The most useful part of the episode is the practical change control model. The recommended flow starts with three clearly defined owners:

  • The requestor, typically IT or the vendor driving the update
  • The operational owner, usually facilities or engineering
  • The incident owner, the person authorized to declare rollback

From there, the team should complete a short pre-checklist before any rollout begins:

  • List the systems that may be impacted
  • Confirm backups or snapshots exist
  • Schedule a short validation window
  • Confirm vendor support availability
  • Pause the rollout if any required item is missing

This structure matters because it removes ambiguity at the exact moment when speed and judgment count most.

The Five Validation Checks to Run in Under 10 Minutes

The guests recommend keeping validation lightweight and fast. Instead of bloated signoff procedures, they suggest five checks that can be run in roughly 10 minutes:

  • Confirm overall system heartbeat
  • Test a basic user workflow such as card access
  • Verify HVAC setpoint response
  • Check one representative telemetry read
  • Perform a sanity check of scheduled jobs

The logic is clear: fast feedback is more useful than long, theoretical coverage if the goal is to catch meaningful operational issues before they spread.

Rollbacks Should Start Small

Rollback planning gets special attention. Full image restores are described as the last resort, not the first step. A better rollback hierarchy starts with the smallest atomic revert possible, such as restoring a configuration file, restarting one service, or reassigning to a failover controller.

Just as important, rollback should be pre-validated on a test unit and documented in a short runbook. When something fails at night, proven steps reduce delay, confusion, and debate.

Another strong recommendation is to keep a canary controller that mirrors production for a period after patching. If the canary remains stable, proceed with wider rollout. If anomalies show up, the team already has logs, evidence, and a defined path to revert.

Communications Matter as Much as the Technical Work

The episode makes the point that patching discipline is also a communications discipline. Tenant-facing notices should explain the timing and possible effects in plain language. Internal operations channels should carry rollback steps, contacts, and escalation guidance.

The suggested communication pattern includes:

  • A 48-hour pre-notice for planned maintenance windows
  • A 2-hour operations alert for pilot rollouts
  • An explicit do-not-escalate clause unless defined triggers occur

That last point is especially practical. It helps prevent minor, expected blips from turning into building-wide alarm and unnecessary wake-ups.

What Success and Failure Actually Look Like

In the success example, a campus staged firmware on a lab rack, ran a 10-minute validation script, patched three controllers as an overnight pilot, monitored a canary for 48 hours, and then completed the weekend rollout with no tenant impact. The differentiators were pre-validation and an explicit rollback owner.

In the failure example, a broad vendor push took place during a vague maintenance window. A thermostat service failed, HVAC dropped into safe mode, and nobody had clear authority to authorize rollback. Tenant calls came in hours later. The lesson is direct: if ownership is unclear and rollback has not been tested, the rollout should not proceed.

Three Actions to Start This Week

The episode closes with three concrete next steps facilities and IT teams can adopt immediately:

  • Document the three-owner model for patch windows
  • Run a one-controller pilot using the five-item validation checklist
  • Create a tiered tenant notification template with explicit escalation triggers

The overall message is practical and disciplined. Patching BAS and OT does not need to be a high-drama overnight gamble. When teams define ownership, keep validation short and meaningful, stage changes carefully, and practice rollback ahead of time, they can improve security without sacrificing reliability.

Deeper dive

Patch Windows for BAS and OT Should Not Feel Like a Midnight Gamble

In this episode of Built, Wired & Secured, the conversation starts with a scenario that immediately frames the real problem. A routine Windows update lands on front-end controllers at 2 a.m. and suddenly half the tenant floors lose HVAC, card access becomes unreliable, and facilities is thrown into crisis mode while leadership tries to figure out who can stop the rollout. That opening sets the tone for the rest of the discussion: patching building automation systems and operational technology is not just an IT maintenance task. It is an operations decision with direct consequences for tenants, staff, and business continuity.

The episode stays intentionally focused on coordination rather than configuration. Instead of going deep into technical tuning, it looks at the policy, ownership, testing discipline, communication steps, and rollback preparation that make patch windows safer and more predictable. The core theme is simple: reliability does not come from hoping a patch behaves. It comes from disciplined maintenance that is practical, repeatable, and clearly owned.

Why BAS and OT Require a Different Update Mindset

One of the clearest points in the episode is that BAS and OT systems cannot be patched like standard user endpoints. In a normal IT environment, many devices are more tolerant of reboots, easier to replace, and less likely to trigger immediate real-world disruption. That is not true in buildings where HVAC schedules, access control, telemetry, and controller dependencies all connect to daily operations.

The discussion breaks this down in operational terms. A reboot that seems minor on paper can interrupt HVAC response, affect access control, interfere with elevator-related logic, or disrupt vendor-connected systems that rely on proprietary integrations. That means the cost of getting patching wrong is not just a failed update. It can quickly become tenant discomfort, front-desk escalations, after-hours vendor calls, and reputation damage.

This is why the episode rejects the idea of a single standard maintenance window for every BAS or OT asset. Different devices carry different risk profiles. Some controllers may be fine to patch after hours. Others should be staged first in a lab environment or piloted on a non-critical floor. The real point is not to create more bureaucracy. It is to match the update approach to the operational risk of the system being touched.

Why Necessary Updates Still Get Deferred

The guests also address a tension most teams recognize: everyone understands the security case for patching, yet updates still get pushed back. The reason, they argue, is not just fear. It is incentives.

Scheduling maintenance windows, coordinating facilities, paying overtime, and involving vendors all have immediate visible costs. By contrast, the risk of a vulnerability being exploited can feel distant and abstract. As a result, patches often slide until they become urgent, and at that point the organization is forced into a less controlled, more disruptive situation.

But the conversation does not stop there. It also pushes back on a different trap: the pursuit of exhaustive testing across every possible variation before doing anything. The guests make a useful distinction between careless testing and prioritized testing. They are not arguing to skip validation. They are arguing to validate the most tenant-critical paths first so progress can happen safely without waiting for a perfect test matrix that never closes.

A Practical Change Control Model for Real Buildings

The strongest operational takeaway from the episode is the lightweight change control model. Instead of heavy process for its own sake, the suggested approach creates enough structure to reduce confusion without slowing the work to a halt.

The model starts by naming three owners before any rollout begins. First is the requestor, usually IT or the vendor initiating the update. Second is the operational owner, typically facilities or engineering, who understands the building-side impact. Third is the incident owner, the person empowered to declare rollback if the update starts causing operational problems.

That third role is especially important because many patch failures become worse when nobody knows who has authority to stop, revert, or escalate. Clear ownership shortens decision time and reduces the chance of avoidable downtime.

From there, the team uses a short pre-checklist:

  • Identify impacted systems
  • Confirm backups or snapshots are available
  • Schedule a defined validation window
  • Verify vendor support availability if needed
  • Pause the rollout if any required item is missing

This is a useful reminder that many failed updates are not caused by the patch itself. They are caused by weak preparation, vague roles, and assumptions about who will handle problems if something goes wrong.

Keep Validation Short, Focused, and Operationally Relevant

Another major point in the episode is the value of a fast validation cycle. The recommendation is not to build a long, complex acceptance plan. It is to run five meaningful checks that can be completed in under 10 minutes and tell the team whether the environment is stable enough to proceed.

Those five checks are:

  • System heartbeat
  • A basic user workflow such as card access
  • HVAC setpoint response
  • One representative telemetry read
  • A sanity check of scheduled jobs

This is effective because it focuses on outcomes that matter in operations. If card access fails, if setpoints are not responding, or if telemetry disappears, the building has a real issue regardless of whether the patch technically installed successfully. Fast feedback based on real workflows is more useful than a long checklist full of items that do not reveal tenant-impacting failure quickly enough.

Why Rollback Preparation Deserves More Attention

The episode treats rollback planning as a first-class part of change control, not a footnote. Full image restores are described as the nuclear option because they are slow, disruptive, and stressful under time pressure. A better approach is to begin with smaller, atomic rollback paths. That might mean restoring a configuration file, restarting a single failed service, or shifting responsibility to a failover controller.

The guests also stress that rollback should be pre-validated on a test unit and documented in a short runbook. That matters because rollback plans often look fine on paper until someone has to execute them under pressure. If the first real test of a rollback procedure happens during an outage, the organization has already lost time it cannot afford.

One especially practical idea is the use of a canary controller that mirrors production for a period after patching. By watching a canary for a week or even 48 hours depending on the environment, teams can surface anomalies early, gather logs, and decide whether a broader rollout is justified.

Communication Is Part of the Control Plan

A building patch window is also a communications event. The episode makes that clear by outlining a tiered notification model. Tenant-facing notices should explain what work is planned, the expected window, and what users may notice. Internal operations notices should include contacts, rollback actions, and escalation guidance.

The recommended timing is practical: 48-hour notice for planned work and a 2-hour operations alert for pilot rollouts. The guests also recommend a do-not-escalate clause unless specific triggers occur. That small addition can prevent expected, low-impact blips from turning into unnecessary alarm across the property.

Just as important, they encourage tabletop drills using those same messages. Who contacts the vendor? Who authorizes rollback? Who updates tenants? If those answers are not clear during a drill, they will not be clear during a real incident.

Success Leaves a Pattern, and Failure Does Too

The success case in the episode is straightforward and repeatable: firmware was staged in a lab rack, validated with a 10-minute script, deployed to three controllers as an overnight pilot, monitored via canary for 48 hours, and then expanded during a weekend rollout. No tenant impact occurred. The reasons are clear: pre-validation, batching, and explicit rollback ownership.

The failure case is just as instructive. A broad vendor push happened during a vague maintenance window. A thermostat service failed, HVAC dropped into safe mode, and there was confusion over who could authorize rollback. Tenant calls came later, and the team lost time because the control structure was weak. The lesson from the episode is blunt and useful: if ownership is unclear and rollback has not been tested, do not proceed.

Three Moves Teams Can Make Right Away

The conversation closes with three immediate actions teams can start this week:

  • Document the three-owner model for patch windows
  • Run a one-controller pilot using the five-item validation checklist
  • Create a tiered tenant notification template with explicit escalation triggers

That is what makes this episode valuable. It does not present patching as an abstract best-practice exercise. It turns patch windows into a manageable operational process: define what breaks, assign authority clearly, validate the workflows tenants actually depend on, and make rollback real before it is needed.

For facilities leaders, IT teams, and building operators, that shift matters. Keeping systems current is important, but so is protecting uptime, tenant confidence, and operational continuity. This episode offers a practical framework for doing both. Listen to the full conversation to hear how disciplined maintenance, short validation cycles, and pre-validated rollback plans can make BAS and OT updates far less risky in the real world.