Show Notes
Why post-incident reviews matter in building technology
Outages and near-misses are part of operating modern buildings, but repeated failures usually have less to do with the initial breakdown and more to do with what happens next. In this episode of Built, Wired & Secured, the conversation centers on a practical question: how do you run a post-incident review that actually leads to change instead of turning into a defensive meeting that goes nowhere?
The discussion opens with a high-impact example: 3 hours, 150 tenants without air conditioning, and a chilled water pump failure on a 95° day. The equipment was repaired, but the real breakdown came afterward. Vendors blamed each other, tenants wanted answers, and the operations team had no clean timeline. That framing sets the tone for the episode: fixing the immediate issue is not enough if the organization does not capture what happened, assign ownership, and make changes that reduce the odds of a repeat.
What a blameless review actually means
A central theme in the episode is the idea of a blameless post-incident review. In plain terms, it is described as a short, timeboxed conversation focused on how systems and decisions combined to produce the outcome. The goal is not to identify who made a mistake. The goal is to understand what failed, what conditions allowed it to fail, and what needs to change.
That distinction matters because many reviews fail before they begin. If the room is defensive, people hold back details, protect their own position, and avoid committing to changes. The guests explain that they deliberately open reviews with a simple expectation: the meeting exists to improve the system, not assign guilt. That single framing helps people provide honest timelines, clarify decision points, and focus on prevention rather than blame.
Where most post-incident reviews break down
The episode identifies several predictable failure points that make reviews ineffective:
- Missing the right people, especially vendors or contractors who made recent changes
- Trying to do deep forensic reconstruction in the first meeting
- Letting the conversation become long, defensive, and unfocused
- Leaving without clear owners and deadlines
- Failing to translate vendor notes or recommendations into tracked internal action items
One example stands out: a failure sequence repeated just three weeks after an earlier incident because everyone assumed the issue was fixed. No one updated the runbook. No one changed the maintenance trigger. And a vendor's written note never became an assigned task in the owner's system. That is the exact gap the episode is trying to solve.
A 30-minute agenda that is simple enough to use
Rather than proposing an elaborate process, the episode lays out a lightweight agenda that can fit into real operating schedules. If you only have 30 minutes, the recommendation is to keep the review to five parts:
- A concise timeline of who did what and when
- The impact, including who and what were affected
- The immediate causes, covering both human and technical actions
- Proposed fixes divided into quick wins and longer life cycle investments
- Owners and deadlines for each action item
The message is clear: if the meeting ends without a name and a date attached to each fix, it did not work. That is presented as non-negotiable.
Preparation also matters. Before the review, the team should gather a short written timeline from facilities and IT along with a few focused evidence sources. The suggested inputs are BAS event log excerpts, vendor change notes, and crew call records. The point is not to build a forensic lab. The point is to collect enough information for the group to agree on what happened and what should change.
Who needs to be in the room
The episode emphasizes that the meeting has to include people who can actually authorize or implement fixes. The must-have attendee list is short but specific:
- The facilities lead
- An IT or network representative
- The vendor or contractor who worked on the system that day
- A tenant operations contact when tenants were affected
If one of those participants cannot attend, the advice is to require a written timeline in advance. Proxies are discouraged unless they know the facts. This matters because a review without someone who can approve a maintenance window, change a BAS sequence, or clarify a vendor action will likely produce discussion without resolution.
How to handle vendor friction without losing momentum
Vendor involvement is treated as a practical challenge, not a theoretical one. Ideally, contracts should require participation. But the episode also acknowledges that contract updates may not happen quickly. The compromise approach is useful: accept written vendor input for the first review, identify a named escalation contact immediately, and schedule a mandatory vendor-required follow-up within 30 days. That keeps the process moving while longer-term contract language is addressed later.
Turning findings into fixes that stick
Another valuable part of the discussion is the framework for prioritizing what comes out of a review. The guests sort recommendations into three buckets:
- Operational fixes, such as process or staffing changes
- Technical fixes, such as BAS sequence adjustments or patches
- Contractual or capital fixes, such as SLA changes or equipment replacement
They recommend prioritizing by risk and effort. Quick operational wins come first, followed by low-effort technical changes, while capital improvements get pushed into budget planning with clear risk notes.
A useful real-world example involves repeated BAS controller reboots after power maintenance. The quick operational win was to change the controller reboot order in the runbook. That update took about 10 minutes to document and test, and the reboot issue stopped. At the same time, the aging controllers were added to a five-year refresh plan. The lesson is that both the immediate low-effort fix and the longer-range capital response matter.
Why recommendations stall
The episode also does a good job of acknowledging why sensible recommendations often die after the meeting. Capital work can stall because leadership wants stronger justification, procurement may dispute pricing, and successful temporary fixes can reduce urgency. Tenant disruption can also complicate replacement decisions. The practical advice is to document those trade-offs instead of letting them disappear. If a recommendation belongs in capital planning, it should include a clear risk note and a target fiscal quarter. If disruption is the reason a fix is deferred, that should be recorded as an accountable choice.
Follow-through is the real test
The simplest discipline in the episode may be the most important. Each action item needs one owner, one deadline, and one place where status is tracked. That could be a shared spreadsheet, a ticket, or a CMMS. The format is less important than consistency. A 30-day follow-up check is also mandatory. If the owner has not updated status before that review, the issue should be escalated.
The guests also recommend capping the post-incident action list at five items. Anything beyond that becomes a project and should move into a different governance path. That keeps the review focused on practical prevention instead of turning into an overloaded wish list.
Key lesson from the episode
The episode ends with a concise and useful takeaway: run the review within 72 hours, keep it blameless, collect minimal evidence, and leave with names, dates, and one quick win that can be implemented within a week. For building operators, facilities leaders, IT teams, and contractors, the message is straightforward. Preventive maintenance is better than emergency repairs, but when an outage does happen, a short and disciplined review is what turns a painful event into a stronger operating system.
After the outage: why the review matters as much as the repair
In building operations, the immediate goal during an outage is obvious: stabilize the situation, restore service, and communicate with affected occupants. But once the equipment is back online, many teams move on too quickly. That is where repeat failures are born.
This episode of Built, Wired & Secured focuses on a discipline that often gets talked about but rarely executed well: the blameless post-incident review. The conversation is grounded in a very practical reality. A chilled water pump goes down on a 95° day. Air conditioning is lost for 150 tenants for three hours. The mechanical issue gets fixed, but the larger failure shows up afterward. Vendors point fingers. Tenants demand explanations. Internal teams scramble to reconstruct a timeline. And three weeks later, the same breakdown sequence appears again because no one truly converted the first incident into lasting operational change.
That is the core problem this episode addresses. In modern building environments, repeating failures are often less about the original fault and more about weak follow-through. If the runbook is not updated, if maintenance triggers are not adjusted, if vendor recommendations never become assigned internal tasks, then the organization has not really learned anything. It has only survived the last outage.
What makes a post-incident review useful
The most important idea in the episode is also the simplest: a useful review is blameless, short, and action-oriented.
Blameless does not mean careless. It means the group is focused on understanding how systems, decisions, timing, communication, and technical conditions combined to create the outcome. That shift matters because blame produces defensiveness, and defensiveness produces incomplete information. If people think the meeting exists to find who messed up, they are less likely to provide honest details about timing, assumptions, missed signals, or unclear ownership.
The discussion offers a practical way to set the tone. Open the meeting with a simple statement that the purpose is to improve the system, not assign guilt. That one sentence can lower the temperature in the room and help the team move faster toward what actually matters: what changed, what was missed, what failed, and what should happen differently next time.
Why many reviews fail before they start
The episode does not present failure as a mystery. It identifies several common patterns that turn reviews into low-value exercises.
First, the wrong people are in the room. If a vendor representative performed the last firmware update and is not included, key context may be missing. If no one present can authorize a BAS sequence change or approve a maintenance window, then the group may talk through the incident without being able to commit to a real fix.
Second, teams often try to do too much in the first meeting. Instead of running a disciplined operational review, they attempt a full forensic reconstruction. That can consume the session, create debate around details that do not affect immediate prevention, and delay obvious corrective actions.
Third, action items are not translated into accountable work. One of the strongest examples in the episode is the vendor note that never became an assigned internal owner. Everyone believed the issue had been addressed, but nothing structural changed. The same failure sequence returned because the organization documented the incident without operationalizing the response.
The 30-minute review template building teams can actually use
A major strength of the conversation is that it does not overcomplicate the process. The proposed agenda is short enough to be realistic for busy operations, facilities, and IT teams.
The review should cover five topics:
- A concise timeline of who did what and when
- The impact on people, spaces, systems, and operations
- The immediate causes, including both human and technical actions
- Proposed fixes divided into quick wins and longer-term life cycle investments
- Named owners and deadlines for each action item
That last point is treated as mandatory. If the meeting ends without a specific person and a specific date attached to each fix, then the review has failed regardless of how thorough the discussion felt.
Just as important, the meeting should happen quickly. The recommendation is to hold it within 72 hours of the outage. That time frame keeps facts fresh, preserves momentum, and reduces the chance that the organization drifts back into normal operations without making changes.
What evidence is enough
Many teams avoid structured reviews because they assume they need a large investigative effort. The episode pushes back on that assumption. For the first review, the team does not need a forensic lab. It needs enough evidence to agree on what happened.
The suggested inputs are straightforward:
- A short written timeline from facilities and IT
- BAS event log excerpts
- Vendor change notes
- Crew call records
This is a useful standard because it balances speed with credibility. It gives the group enough documentation to prevent memory-driven disagreements without slowing the review down with exhaustive analysis. When the goal is short-term learning and prevention, a coherent and shared account of the event is more valuable than a perfect reconstruction that arrives too late.
Who should attend and why that matters
The attendee list reflects the reality that building technology incidents cross disciplines. Facilities may own the physical plant. IT may support network dependencies. Vendors may control service history, firmware, or specialized systems. Tenant operations may understand the real-world impact of a disruption better than anyone else.
That is why the episode recommends four must-have participants:
- The facilities lead
- An IT or network representative
- The vendor or contractor involved that day
- A tenant operations contact when tenants were affected
If one of those people cannot attend, the fallback is to require a written timeline in advance. The key principle is not attendance for its own sake. It is decision-making capacity. A review is only useful if the people present can supply the facts and support the resulting changes.
Managing vendor accountability without waiting for perfect contracts
Vendor participation can be one of the hardest parts of any post-incident process. The episode acknowledges that while better contract language may be necessary, contract revisions are often slow. In the meantime, teams still need a workable process.
The recommended compromise is smart. Require written vendor input right away. Identify a named escalation contact. Then schedule a mandatory vendor follow-up within 30 days. That approach avoids losing momentum while still building pressure for stronger contractual accountability in the next cycle.
For property owners and operators, this is an important business lesson. Better building resilience is not just a technical issue. It is also a governance issue. If vendors can influence system outcomes but are not required to participate in learning from failures, the owner is carrying operational risk without enough leverage.
Sorting fixes into the right buckets
Another practical takeaway is the three-bucket framework for post-incident recommendations:
- Operational fixes, such as process changes or staffing adjustments
- Technical fixes, such as BAS sequence updates or patches
- Contractual or capital fixes, such as SLA changes or equipment replacement
This matters because not every issue should be treated the same way. Some risks can be reduced immediately with a runbook update or a maintenance change. Others require budget planning or contract negotiation. By separating these categories, teams can move quickly on what is easy while still preserving visibility into what is larger and slower.
The example of repeated BAS controller reboots after power maintenance illustrates this perfectly. The quick operational fix was to change the reboot order in the runbook and test it. That stopped the immediate issue. But the controllers were also aging, so they were added to a five-year refresh plan. In other words, the team handled both the short-term operational weakness and the long-term asset risk.
Why capital recommendations disappear
The episode is especially useful in how honestly it addresses stalled recommendations. Capital requests do not fail only because teams forgot about them. They often stall because leadership wants stronger justification, procurement challenges pricing, or temporary improvements reduce the apparent urgency. Tenant disruption can also change the equation, especially when replacement work affects occupied spaces.
The right response is not frustration alone. It is documentation. If a replacement is deferred, the risk should be written down. If the decision is to avoid disruption now and accept higher failure exposure later, that should be recorded as a conscious trade-off. If capital work is needed, it should be assigned to a target fiscal quarter rather than left as a vague future intention.
That level of documentation turns an operational memory into an organizational decision. It also gives leadership a clearer basis for future planning.
The follow-through discipline that prevents repeat outages
Ultimately, the episode argues that the real success or failure of a review is decided after the meeting. Two simple rules stand out.
First, every action must have one owner and one deadline. Shared ownership often means no ownership. Second, the team must schedule a 30-day check to review progress and escalate stale items. Tracking can live in a shared spreadsheet, a ticket, or a CMMS. The platform matters less than the consistency.
The guests also recommend capping the review's action list at five items. That is an important discipline. Too many action items create diffusion, weak follow-through, and clutter. If the list is longer than five, some of the work probably belongs in a formal project or preventive maintenance program instead of a post-incident action tracker.
A simple process with outsized value
The biggest takeaway from this episode is that resilient building operations do not always require a heavy process. They require a repeatable one. A 30-minute blameless review, held within 72 hours, supported by a short timeline and a few evidence sources, can produce meaningful change if the team includes the right people and leaves with specific owners and dates.
For property owners, facilities teams, IT leaders, and contractors, this is where operational discipline turns into business value. Better reviews improve uptime, support clearer tenant communication, strengthen vendor accountability, and help capital planning reflect real risk instead of assumptions.
If you want a practical way to start, the episode points listeners to a one-page checklist and reinforces one final habit worth adopting immediately: schedule the 30-day follow-up at the same time you schedule the first review. That one calendar decision can be the difference between temporary relief and lasting improvement.
If this topic is relevant to your building or portfolio, this episode is worth a listen because it keeps the process simple, accountable, and grounded in what teams can actually do after the next outage.