Show Notes
Why coordination failures create the worst building outages
This episode of Built, Wired, and Secured focuses on a problem that shows up in real buildings every day: many disruptions are not caused by a catastrophic hardware failure, but by a breakdown in ownership, communication, and handoffs between teams. The opening scenario makes that risk tangible. During a daytime transfer, a maintenance tech flips a breaker. Nothing appears to fail immediately, but the backup power handoff never completes. Elevators stop between floors, cooling begins to drift, and the front desk gets overwhelmed with calls. Two hours later, the team discovers that no one on the call actually owned the ATS logic.
That example frames the core question of the episode: could a short, focused tabletop have exposed the missing handoff before occupants were affected? The answer from the guests is yes. Tabletop exercises are presented here not as large-scale, expensive events, but as practical, time-boxed working sessions that help facilities, IT, security, and vendors identify weak points before an outage forces everyone into reactive mode.
What a realistic tabletop should look like
One of the most useful takeaways in the conversation is that an effective tabletop does not need to be large or elaborate. The recommended format is simple:
- Choose one scenario only
- Limit the session to 60 to 90 minutes
- Set one primary objective tied to business impact
- Use realistic constraints such as limited staff or no extra budget
The guests stress that scope discipline matters. If teams try to solve every possible problem in one meeting, they end up solving none of them. A better approach is to pick one event, such as a utility-to-generator transfer, a BAS cloud failure, or a carrier edge outage, and define success in measurable terms. Examples from the episode include restoring emergency lighting within 20 minutes, confirming who is responsible for calling the carrier, or verifying that manual elevator recall can be performed.
This emphasis on measurable outcomes is important because it connects technical readiness to operational and business consequences. The discussion repeatedly brings the focus back to tenant impact, revenue sensitivity, and service continuity.
How to design a scenario leadership will care about
The guests recommend starting scenario design with a simple question: what breaks if this goes down? That framing keeps the exercise grounded in consequences people care about. Elevators, cooling, and emergency lighting are highlighted as examples because they directly affect tenants and building operations.
From there, the exercise should identify the actual dependency source. Rather than using a vague scenario like “power outage,” teams should define the specific trigger, such as scheduled maintenance, utility transfer, or UPS failure. The reason is that each trigger creates different decision points, ownership questions, and remediation paths.
That level of clarity makes it easier for leadership to see value. Instead of participating in a generic technical discussion, leaders can evaluate a concrete question: if this exact failure occurs under these constraints, do we know who acts, in what order, and how quickly?
Who needs to be in the room
The attendee list recommended in the episode is intentionally practical. The right group includes:
- Building engineering
- IT network leadership
- Security
- The primary vendor for the relevant system, such as generator or BAS
- A property manager
The group also recommends using a neutral moderator. That role is not there to provide deep technical expertise, but to keep the meeting moving, hold the clock, and prevent the conversation from devolving into an endless technical workshop. The suggested agenda is tight:
- 5 minutes for scenario and objective
- 20 minutes to confirm handoffs
- The remainder to play decisions and capture outcomes
This structure reinforces the episode’s main theme: the goal is not theoretical perfection. The goal is to expose practical gaps quickly enough that teams can fix them.
The minimum artifacts that drive real change
A standout part of the conversation is the focus on lightweight documentation. Rather than creating long reports that no one will read, the guests recommend producing just three short artifacts during or immediately after the session:
- A one-page owner matrix listing components and accountable contacts
- A short playbook snippet with decision triggers and two to three manual workaround steps
- A playback note capturing decisions, owners, and deadlines
Keeping each artifact to one page is intentional. The point is usability. In a real event, teams need to know who owns what, what manual fallback exists, and what actions were agreed to. Short documents increase the odds that the material will actually be referenced and maintained.
The follow-up discipline matters just as much. The guests advise assigning owners in the room and scheduling a 30-day checkpoint before the meeting ends. Their point is blunt and accurate: if the follow-up is not scheduled, it usually does not happen.
What these drills tend to reveal
The most common surprises are not especially glamorous, which is exactly why they are so dangerous. The episode highlights several recurring findings:
- Unowned components, such as ATS logic, surge protection at tenant risers, or third-party UPS systems
- Vendor responsibility gaps where ownership effectively stops at a connector or demarcation point
- Communication assumptions, especially around who has the carrier contact information and authority to act
These are the kinds of issues that can turn a manageable incident into a prolonged outage. The discussion makes clear that many teams assume these details are already understood until a drill forces them to answer the question in real time.
Examples of measurable wins
The episode closes the loop with concrete examples. In one 90-minute transfer drill, the team learned that the generator would start, but ATS logic in the BAS had effectively been delegated to a vendor who only worked weekdays. The fix was straightforward: update the owner matrix, add a manual override to the playbook, and schedule a firmware check. The result was meaningful: recovery time dropped by an hour the next time the issue came up.
In another example, a tabletop revealed that cooling fallback would not run when the BAS cloud entered read-only mode. The remediation was similarly practical: document a manual cooling sequence, train two technicians, and push the task into the maintenance backlog. That preparation later prevented a service level breach.
The five-step checklist listeners can use now
The episode ends with a clear checklist:
- Pick one realistic scenario tied to tenant impact
- Limit the session to 60 to 90 minutes and time-box the agenda
- Invite operational owners and one critical vendor representative
- Capture an owner matrix, a playbook snippet, and a playback note during the session
- Assign owners, deadlines, and a 30-day follow-up
The final advice is to start small and repeat often. The guests also recommend measuring time to identify the owner, time to implement a workaround, and whether that workaround prevented tenant impact. Those metrics help prove value to leadership and support a repeatable resilience program instead of a one-off exercise.
Why building outages often start with coordination failures, not broken hardware
When people think about building outages, they often picture a major device failure: a dead generator, a failed UPS, or a network core going dark. This episode of Built, Wired, and Secured argues that the real problem is often more basic and more dangerous. The outage is extended not because a component failed, but because nobody is fully clear on ownership, handoffs, manual workarounds, or who has authority to act.
The opening example makes the point immediately. A maintenance tech flips a breaker during a daytime transfer. At first, nothing seems obviously wrong. Then the backup power handoff never completes. Elevators stop between floors. Cooling begins to drift. The front desk is suddenly handling a flood of calls. Hours later, the team discovers that the ATS logic was not clearly owned by anyone participating in the response.
That is the kind of failure a tabletop exercise is designed to catch before a real incident happens. In this conversation, Alex Morgan speaks with Michael Harrington and James Rogers about how short, realistic drills can expose hidden dependencies and produce improvements that facilities and IT teams can actually implement.
Tabletops do not need to be large to be valuable
One of the most useful myths this episode clears up is the idea that a tabletop exercise has to be a major undertaking. The format recommended here is intentionally lean. Pick one scenario. Set one primary objective. Keep the session to 60 to 90 minutes. Make the outcome measurable.
That narrow scope is what makes the exercise useful. A team might focus on a utility-to-generator transfer, a BAS cloud failure, or a carrier edge outage. What matters is not the complexity of the event but the clarity of the goal. If a session tries to address every possible outage path, it usually produces vague discussion and little follow-through. If it focuses on one realistic situation, it can surface the exact handoffs and decisions that matter.
The guests repeatedly connect this discipline back to business impact. Technical exercises become more credible when the objective is phrased in operational terms: restore emergency lighting within 20 minutes, verify manual elevator recall, or confirm who is responsible for calling the carrier during an outage. That shift matters because leadership does not fund resilience for its own sake. Leadership funds reduced tenant disruption, reduced downtime, and fewer preventable escalations.
Start scenario design with impact, not infrastructure
A practical lesson from the episode is how to choose the right scenario. The recommended first question is simple: what breaks if this goes down? That question changes the conversation from infrastructure inventory to operational consequence.
If elevators, cooling, and emergency lighting are the first systems affected, then the scenario should be built around those outcomes. The team can then define one success criterion tied directly to the risk. Instead of saying, “We want to review backup power,” the group might say, “We want to confirm whether a manual workaround can maintain safe elevator and lighting operations during a failed transfer.”
The guests also warn against vague scenario definitions. “Power outage” is too broad to be useful. Teams should identify the dependency source precisely: utility transfer, scheduled maintenance, or UPS failure. Each creates different ownership questions, different communication paths, and different remediation options. That specificity is what turns a general discussion into a working drill.
The right people matter more than a large audience
Another strong point in the discussion is that attendance should be selective, not exhaustive. The goal is not to gather every possible stakeholder. The goal is to have the people who actually own decisions and systems in the room.
The recommended attendee list includes building engineering, an IT network lead, security, a property manager, and the primary vendor relevant to the scenario, such as the generator or BAS vendor. This cross-functional mix reflects the reality of modern building operations. Power, controls, telecom, security, and tenant experience are closely linked, even when ownership is distributed across separate teams or companies.
The episode also recommends a neutral moderator. That role is critical because technical teams often slip into troubleshooting mode too early. A neutral facilitator keeps the group on time, prevents deep dives from consuming the session, and ensures the meeting produces decisions rather than just discussion.
The suggested agenda is refreshingly practical:
- Use the first five minutes to state the scenario and objective
- Use the next twenty minutes to confirm handoffs and responsibilities
- Use the remaining time to walk decisions, workarounds, and outcomes
This keeps the exercise focused on accountability and execution rather than theory.
Three artifacts that turn discussion into action
Many resilience efforts fail because the documentation they produce is too long, too abstract, or too disconnected from day-to-day operations. This episode argues for the opposite approach: short documents with immediate operational value.
The three recommended artifacts are simple:
- A one-page owner matrix showing components and accountable contacts
- A short playbook snippet with decision triggers and manual workaround steps
- A playback note listing decisions, owners, and deadlines
These are not meant to become shelfware. They are meant to help people act. In a real event, teams do not need a lengthy postmortem template. They need to know who owns the ATS logic, who can call the carrier, what manual cooling sequence is available, and what commitments were made to close the gaps discovered in the drill.
The conversation also stresses the importance of scheduling follow-up before the meeting ends. A 30-day checkpoint gives the exercise teeth. Without that checkpoint, the owner matrix goes stale, the workaround never gets documented, and the unresolved issue remains in place until the next outage exposes it under pressure.
The most common findings are the ones teams assume are already covered
Perhaps the most valuable insight in the episode is what these drills tend to uncover. The surprises are usually not exotic failures. They are ordinary gaps in ownership and communication that have simply never been tested.
The guests mention unowned components such as surge protection at tenant risers, third-party UPS systems, and ATS logic. They also describe vendor boundary issues where responsibility effectively ends at a connector, leaving no one clearly accountable for the system behavior that matters most to occupants. On the communication side, teams often assume someone has the carrier contact information and authorization to engage support. A tabletop makes that assumption visible very quickly.
These are exactly the kinds of overlooked details that extend downtime. A building may have the right hardware in place and still fail operationally if nobody can answer basic response questions in real time.
Examples of remediation that actually improved outcomes
The episode stands out because it does not stop at theory. It gives concrete examples of what changed after a drill.
In one 90-minute transfer exercise, the team discovered that the generator would start, but ATS logic in the BAS depended on a vendor who only worked weekdays. The remediation was not massive or expensive. The team updated the owner matrix, added a manual override to the playbook, and scheduled a firmware check. The practical result was significant: when the issue occurred again, time to recover dropped by an hour.
In another case, a tabletop revealed that cooling fallback would fail if the BAS cloud went read-only. The team documented a manual cooling sequence, trained two technicians, and fed the related work into the maintenance backlog. Later, that training helped prevent a service level breach.
These examples reinforce the larger message of the episode: resilience often improves through clear ownership, small documentation updates, focused training, and disciplined follow-up.
A simple checklist to start this week
The closing checklist is one of the most actionable parts of the conversation:
- Pick one realistic scenario tied to tenant impact
- Limit the session to 60 to 90 minutes and time-box the agenda
- Invite operational owners and one critical vendor representative
- Capture an owner matrix, a playbook snippet, and a playback note during the session
- Assign owners, deadlines, and a 30-day follow-up
The guests also recommend tracking three data points during early drills: time to identify the owner, time to implement a workaround, and whether that workaround prevented tenant impact. Those simple metrics help leadership see the value of continued investment in the program.
If there is one idea to take from this episode, it is that preventive maintenance is not only about inspecting equipment. It is also about testing human coordination around the systems that matter most. A short tabletop can reveal hidden handoffs before they become visible to tenants, occupants, and leadership in the middle of an outage.
If you want a practical starting point, the episode points listeners to a one-page tabletop checklist at gdstechnology.io/bws. It is a fitting extension of the conversation: keep the exercise simple, make it measurable, and turn what you learn into actions that reduce real operational risk.