Show Notes
Why Resilience Drills Matter
This episode of Built, Wired & Secured focuses on a problem many building operators know too well: a small maintenance action can trigger a much larger operational failure when hidden dependencies are not fully understood. The opening scenario is simple but familiar. A vendor flips what they believe is a maintenance breaker, and suddenly an access control panel and a network closet lose power. Within half an hour, an entire floor cannot badge in, alarms are firing in the BAS environment, and reception is flooded by tenants looking for answers.
That setup frames the core issue of the episode. Many organizations believe they are prepared because they run tabletop exercises, but those exercises often stay too high-level to reveal the real weak points. The discussion makes the case that resilience drills need to move beyond paperwork and into controlled, measurable testing that validates both technology and operations.
Why Tabletop Exercises Fall Short
The conversation highlights two reasons teams rely too heavily on tabletop exercises: convenience and false confidence. Tabletop sessions are inexpensive, low-risk, and easy to organize, so they become the default. But the tradeoff is that they rarely test the brittle edges where actual failures occur.
- They do not expose undocumented dependencies between systems.
- They often skip real vendor handoffs and remote access challenges.
- They can overlook delays caused by MFA or contact escalation issues.
- They may assume that backup power and failover processes work without verifying the actual load paths.
One of the clearest insights in the episode is that drills must test not only written procedures, but also the real wiring between BAS, UPS, access control, and network infrastructure. That is where assumptions tend to break under pressure.
Start with a Clear Objective
Rather than trying to test everything at once, the episode recommends starting with a narrow objective. Before designing a drill, teams should ask what behavior they are actually trying to validate. That might mean testing vendor coordination, cross-system failover, tenant continuity, or staff response times. The key is to choose one or two priorities and keep the scope tight.
Instead of testing an entire campus, for example, the drill might focus on a single BAS controller and its network path. That smaller scope makes the exercise safer, easier to control, and more likely to produce useful results.
The speakers also stress the importance of defining measurable success criteria in advance. “It worked” is not enough. Teams need targets that can be evaluated after the fact.
- How many minutes did it take to recover the controller?
- How quickly was remote vendor access restored?
- How long did it take to notify tenants?
- How many unexpected dependencies were identified?
Balancing Realism and Safety
A major theme in the episode is the tension between realistic testing and operational risk. Live tests reveal the problems that simulations miss, but live tests also create the possibility of disruption. The recommended answer is a tiered testing model.
- First, simulate the scenario in a lab environment.
- Next, validate the process in an isolated zone.
- Then, run a targeted live test during a low-impact window.
Every live step should include a rollback plan and a clearly defined emergency stop that any technician can invoke. That approach makes testing more practical and reduces the chance that a drill creates the very outage it is meant to prevent.
The discussion also ties resilience drills to preventive maintenance. When drills are part of the PM cadence, teams become more familiar with the process, surprises decrease, and lessons learned are easier to turn into ongoing operational improvements.
A Practical BAS Drill Example
One of the most useful parts of the episode is the staged BAS example. Instead of jumping straight into a disruptive live test, the speakers describe a sequence designed to surface issues gradually.
- Step one: simulate sensor loss in the lab by emulating the sensor feed so the controller behaves as if the sensor has failed.
- Step two: force a controller swap in an isolated rooftop unit during overnight hours.
- Step three: bring an adjacent zone online under supervision.
Each stage has a single owner and a rollback plan. That structure creates accountability and keeps the test controlled. In the example discussed, this type of staged live validation exposed a vendor restart script that did not fully restore service. A tabletop exercise would not have found that issue.
Tenant Coordination and Staffing
The episode makes it clear that resilience is not only a technical issue. Tenant communication and staffing structure are just as important. From a tenant coordination standpoint, the speakers call out several non-negotiables:
- A short pre-test briefing.
- Written consent for critical tenants.
- A contact tree with escalation times.
- Contingency services in place before live testing begins.
One example from the conversation stands out: a vendor’s MFA requirement delayed remote access and cost the team 20 minutes. An alternate contact prevented the delay from becoming worse. It is a good reminder that “small” process gaps become major operational problems during a time-sensitive event.
On staffing, the recommendation is to keep the on-site team lean but effective. A strong model includes two technical leads, a facilities manager, and a tenant liaison, while remote command includes vendor leads and network support. Too many people on site create confusion. Too few create blind spots.
Three Actions to Take This Quarter
The episode closes with practical direction for teams that want to get started now.
- Build and maintain an asset-driven scenario list based on real dependencies such as UPS feeds, network paths, and vendor access.
- Define measurable success criteria and communication protocols before testing begins.
- Use a tiered model: simulate, isolate, and then run targeted live drills with rollback plans and tenant notice.
Recommended metrics include time to detect, time to recover, vendor response time, number of escalations to business continuity, and unknown dependencies found.
Finally, every drill should end with a blameless review. Capture what surprised the team, update SOPs, and move any newly identified preventive maintenance tasks into the maintenance schedule. Vendor commitments should also be documented in writing, including response windows, escalation contacts, and approved remote access methods.
The Bottom Line
The central message of this episode is simple: resilience drills should answer a direct operational question, not just satisfy a compliance checkbox. If a system goes down, what breaks next? The only way to answer that with confidence is through realistic, scoped, well-coordinated testing.
For building owners and operators, the value is not in running bigger drills. It is in running smarter ones that reveal hidden dependencies, strengthen communication, and turn surprise into practice before a real incident forces the issue.
Resilience Drills for Modern Buildings: How to Test Before Failure Tests You
In modern buildings, operational resilience is rarely determined by a single system. It is determined by how systems interact under stress. Access control depends on power. BAS workflows depend on controllers, network paths, and vendor access. Tenant experience depends on how quickly staff can detect, communicate, and recover from disruption. When one overlooked dependency fails, the result can escalate quickly.
That is the core lesson from this episode of Built, Wired & Secured, which explores how building owners and operators can design resilience drills that are realistic, safe, and useful. The discussion moves past theory and into practical testing: what to test, how to stage it, how to communicate with tenants and vendors, and how to measure whether the drill actually improved readiness.
Why Small Failures Become Building-Wide Problems
The episode opens with a vivid example. A vendor flips what seems like a maintenance breaker. Instead of a routine task, the action cuts power to both an access control panel and a network closet. Within 30 minutes, a floor can no longer badge in, BAS alarms are firing, and tenants are pressing for answers.
That kind of cascade is exactly what resilience drills are meant to prevent. The original failure is not always the real issue. The larger problem is often the undocumented dependency that nobody accounted for: a shared feed, an untested failover path, a remote access requirement, or a restart process that only works on paper.
For building operations teams, this is where resilience becomes more than a technical exercise. It becomes a business continuity issue. Tenants feel disruption immediately, and operations teams are forced to respond in real time under pressure.
The Problem with High-Level Tabletop Exercises
The conversation makes a strong case that tabletop exercises, while useful, are not enough on their own. Teams default to tabletop because it is cheap, safe, and easy to run. The problem is that tabletop testing often validates assumptions rather than systems.
In a tabletop session, it is easy to say, “We lose power, we fail over.” It is much harder to validate whether the UPS is actually carrying the right loads, whether a vendor can log in remotely without delay, or whether the access control and BAS environments have an undocumented dependency on the same circuit or network segment.
That gap between procedural confidence and operational reality is where risk hides. Teams may leave a tabletop session feeling prepared while never having tested the actual handoffs between facilities, network support, vendors, and tenant-facing staff.
The episode emphasizes that the most valuable drills are the ones that exercise the handoffs. In other words, resilience is not just about whether the equipment can recover. It is also about whether the people, vendors, communications, and escalation paths work when timing matters.
Start with a Narrow, Measurable Goal
One of the most practical recommendations in the episode is to begin every drill with a tightly defined objective. Instead of asking, “Can we handle a failure?” teams should ask, “What behavior are we validating?”
That behavior might involve:
- Vendor coordination during a controller failure
- Cross-system failover between BAS and supporting infrastructure
- Tenant continuity during a localized outage
- Staff response times and communication accuracy
Keeping scope tight is critical. Rather than testing an entire facility, a team might focus on a single BAS controller and its network path. That narrower scope reduces risk and produces findings that are easier to act on.
Equally important, success criteria should be defined before the drill starts. “It worked” is too vague to support improvement. Better targets include measurable restoration times, communication deadlines, escalation thresholds, and documented recovery steps.
If a team cannot define success before the test, it will struggle to learn from the results afterward.
Use a Tiered Testing Model
The episode outlines a practical middle ground between safe simulation and risky live testing. Rather than choosing one or the other, teams should tier their drills.
- Simulate the scenario in the lab first.
- Validate in isolated zones second.
- Run targeted live tests last, during low-impact windows.
This structure matters because simulation alone can hide real-world issues, while live testing without preparation can create unnecessary exposure. A tiered approach lets teams validate logic, process, and human coordination before they touch live environments.
The speakers are clear that live tests are necessary for critical paths. If the goal is to know whether a failure process really works, at some point it must be exercised under live conditions. But those tests need controls: limited zones, defined time boxes, tenant coordination, contingency services, rollback procedures, and an emergency stop that any technician can activate.
That is what turns live testing from a gamble into a disciplined operational practice.
A Strong Example: BAS Failover in Stages
The BAS example in the episode provides a useful template for building teams. Rather than performing a broad or disruptive test, the sequence is staged to reveal problems progressively.
First, the team simulates sensor loss in the lab by emulating the feed to the controller. This allows logic and alarm behavior to be tested without changing live HVAC operation.
Second, the team performs a controller swap in an isolated rooftop unit during overnight hours. This introduces live conditions while limiting operational exposure.
Third, an adjacent zone is brought online under supervision. That step confirms whether the environment behaves correctly as scope expands.
Each stage has a single owner and a rollback plan. That level of discipline is a recurring theme in the episode. Ownership, reversibility, and scope control are what make drills safe enough to run and useful enough to trust.
Importantly, this staged live process uncovered a vendor restart script that did not fully restore service. That kind of defect would likely have gone unnoticed in a tabletop review. The finding illustrates the broader point of the episode: realistic drills expose hidden weaknesses that process reviews alone cannot catch.
Tenant Communication Is Part of Resilience
Operational resilience is not only about uptime. It is also about trust. The discussion makes the case that tenant coordination should be built into every resilience drill, especially when live testing is involved.
The recommended basics are straightforward:
- Provide a short pre-test briefing.
- Get written consent from critical tenants.
- Build and verify a contact tree with escalation times.
- Ensure contingency services are available if the test impacts operations.
These steps reduce confusion and make it easier to manage expectations if something unexpected occurs. They also reinforce that resilience planning is a service issue, not just an engineering issue.
The episode gives a good example of how a minor process gap can create a major delay: a vendor’s MFA blocked remote access, costing 20 minutes. An alternate contact saved the situation. The lesson is simple. Response plans should account for real constraints like authentication, scheduling, and contact availability, not just equipment status.
Staffing the Drill Without Creating Chaos
The conversation also addresses staffing. Too many people on site can slow decisions and create conflicting instructions. Too few can leave critical blind spots uncovered.
The model recommended in the episode is lean and intentional: two technical leads, a facilities manager, and a tenant liaison on site, with vendor leads and network support participating remotely. That structure keeps the on-site team decisive while still ensuring the right expertise is available.
Communication protocol matters just as much as headcount. Teams should know who speaks to tenants, who escalates issues, who coordinates vendors, and who calls for rollback if the test begins to drift outside acceptable risk.
In practice, drills are not just a way to test systems. They are also a way to test command structure.
The Metrics That Matter
A useful drill generates more than anecdotes. It produces operational data. The episode recommends tracking:
- Time to detect
- Time to recover
- Vendor response time
- Number of escalations to business continuity
- Unknown dependencies found
That last metric may be the most valuable. Hidden dependencies are often painful to discover, but finding them during a drill is far better than finding them during a real outage.
After every exercise, the team should conduct a blameless review. Capture what happened, document surprises, update SOPs, and move any newly identified PM tasks directly into the maintenance schedule. Vendor expectations should also be formalized in writing, including drill participation, response windows, escalation contacts, and approved remote access methods.
From Compliance Exercise to Operational Discipline
The most important takeaway from this episode is that resilience drills should not be polite. They should answer the direct question: what breaks if this goes down?
That mindset changes the purpose of the exercise. Instead of checking a box, the team is trying to expose risk early enough to fix it. That means narrowing scope, measuring outcomes, validating handoffs, and testing real dependencies in controlled ways.
For building operators, facilities leaders, and technology stakeholders, this is where resilience becomes practical. The goal is not to create more disruption. It is to create more confidence by replacing assumptions with evidence.
If your organization has been relying on tabletop sessions alone, this episode offers a better path forward: scenario-driven planning, tiered testing, tenant-aware communication, and post-drill process improvement. That is how surprise becomes practice.
To hear the full conversation and the examples behind these recommendations, listen to the complete episode of Built, Wired & Secured.