Show Notes
When a Small Hardware Failure Becomes a Big Building Problem
A failed BAS controller at 3:00 a.m. can trigger a chain reaction that reaches far beyond the mechanical room. In this episode of Built, Wired & Secured, Alex Morgan sits down with Michael Harrington and James Rogers to outline a practical spare-parts playbook for commercial buildings. The focus is simple: identify the few components that can stop operations cold, decide what should be owned versus sourced quickly, and make sure stored spares are actually usable when the moment comes.
The conversation starts with a realistic scenario. A BAS controller fails overnight. There is no spare on site. Elevators get warm, conference rooms are evacuated, tenants begin calling, and what should have been a short repair stretches into four days. The takeaway is immediate: spare-parts planning is not about having shelves full of inventory. It is about protecting business continuity by owning and maintaining the few parts that truly matter.
Start With Impact, Not Part Numbers
Michael recommends beginning with impact rather than a parts catalog. Instead of asking what components exist in the building, teams should ask which failures create immediate tenant disruption or safety risk. He points to examples such as UPS modules, network uplink modules, main BAS controllers, and door access controllers.
From there, the real question becomes recovery time. If a missing part would extend recovery beyond a shift, it moves higher on the priority list. James adds that failure frequency matters too, but not in isolation. A part that rarely fails can still be highly critical if replacement takes weeks. That is why both guests support a simple criticality matrix built around consequence and time to recover.
- High consequence plus long replacement time means immediate action
- Low consequence with fast replacement may not justify on-site stock
- Rare failures still matter when procurement lead times are long
- The matrix should guide decisions, not become an inventory wish list
When to Stock and When to Source
One of the strongest themes in the episode is that stocking parts and relying on vendors are both trade-offs. James is skeptical of depending entirely on vendor promises, especially when holidays, factory delays, or logistics issues can wipe out an SLA in practice. Michael acknowledges that owning spares ties up capital and introduces obsolescence risk, but he argues that true single points of failure deserve a different standard.
The group lands on a practical rule: own spares for parts where vendor lead times exceed your outage tolerance. Lower-risk items can be sourced, but only if procurement performance is tested and documented. Michael describes a straightforward way to validate vendor claims: run a mock order, time it, verify delivery windows, and make sure there is a clear escalation path and backup source if the vendor misses.
Obsolescence and Storage Can Ruin a Good Plan
Simply buying a spare is not enough. The episode highlights how quickly inventory can become unreliable if it is left unmanaged. Michael shares a case where a new controller from inventory did not match the live network firmware, costing hours in reconciliation. That experience reinforces the need to track firmware, board revisions, and compatibility right alongside the physical part.
James brings the issue down to day-to-day maintenance reality: storage conditions matter. Humidity, temperature, batteries, and packaging all affect shelf life. A spare that sits untouched in the wrong environment can become clutter instead of protection.
- Track firmware versions and hardware revisions in the spare record
- Update, repurpose, or retire spares that fall outside compatibility windows
- Store boards in anti-static, dry conditions
- Label every spare with a test-by or use-by date
- Set quarterly boot and test schedules for critical components
Use CMMS to Make Spares Operational
James makes a strong case for treating spares as first-class records inside the CMMS. Each spare should have an ID, a storage location, a last test date, a firmware version, and a scheduled test interval. This turns spare management into a repeatable operating process rather than an informal habit.
He also recommends rotating spares into low-risk circuits during preventive maintenance. That approach does two things at once: it exercises parts before an emergency exposes a hidden problem, and it gives technicians hands-on replacement practice. The operational benefit is clear. When an outage happens, the team is not learning in real time.
Cross-training is another practical point. A spare does not help much if only one technician knows where it is stored or how to install it. Teaching at least one colleague how to locate and swap the part is described as cheap insurance.
Two Real Cases That Show the Difference
The episode includes two concise examples that capture the entire argument. In the first, a campus kept a spare network uplink module in a locked cabinet. When a contractor accidentally fried a live port at 2:00 a.m., the team hot-swapped the module, triggered failover, and restored connectivity in 20 minutes with no tenant disruption. In that case, the spare paid for itself almost immediately.
In the second, an aging BAS spare stored in a damp utility closet failed to boot when it was finally needed. Corroded connectors turned a supposed backup into a useless liability, and the building lost HVAC control for days. The team later improved storage, added quarterly boot tests, and properly logged the spare in CMMS, but only after learning the lesson the hard way.
Three Rules to Apply This Week
Before closing, the group leaves listeners with three clear actions:
- Build a one-page criticality matrix for your top building systems
- Put spares into your CMMS with IDs, test intervals, and rotation schedules
- Run a live vendor SLA test annually and confirm the backup plan if delivery fails
They also suggest a simple 15-minute audit for one system this month: UPS, network, BAS, or access control. Identify the vulnerable part, confirm who supplies it, and determine whether a tested spare exists. If it does not, document it in CMMS or add it to the next capital plan.
Why This Matters
The core message of this episode is that spare-parts planning is really continuity planning. The right spare, in the right condition, with the right documentation, can prevent a minor equipment failure from turning into tenant disruption, overtime costs, and emergency decision-making. Good operations often go unnoticed, and that is exactly the point. The best day in building operations is the one nobody notices.
If your team has never formalized how it identifies, stores, tests, and justifies critical spares, this episode offers a practical framework to start now without overcomplicating the process.
A Spare Parts Strategy Is Really a Continuity Strategy
In commercial buildings, hardware failures rarely stay confined to the hardware itself. A failed controller, uplink module, or access device can quickly become a tenant experience problem, an operations problem, and a finance problem all at once. In this episode of Built, Wired & Secured, Alex Morgan talks with Michael Harrington and James Rogers about a practical spare-parts playbook for keeping building systems running when hardware fails.
The discussion is grounded in a scenario that feels familiar to anyone responsible for facilities, operations, or building technology. A BAS controller fails at 3:00 a.m. There is no spare on site. Elevators get warm, conference rooms are evacuated, tenants begin calling the front desk, and what should have been a short repair stretches into four days of disruption, overtime, and refund pressure. The lesson is not that every building needs a warehouse full of inventory. It is that teams need a disciplined way to identify and maintain the few spares that truly protect operations.
Start With Operational Impact
One of the most useful ideas in the conversation is Michael’s advice to start with impact, not part numbers. Many organizations approach spare planning like a procurement exercise. They begin by reviewing models, SKUs, or supplier catalogs. That may be necessary later, but it is not the right first move. The right first move is to ask which failures would immediately disrupt tenants or create safety risk.
That shift in thinking changes the conversation. Instead of getting lost in a long parts list, teams can focus on systems that have a direct operational consequence when they fail. In the episode, examples include UPS modules, network uplink modules, main BAS controllers, and door access controllers. These are not just components. They are points where hardware reliability meets occupant experience.
Michael then adds a second filter: how long would it take to recover if the part were unavailable? If the answer is more than a shift, the part becomes more strategically important. James reinforces the point by noting that failure frequency matters, but only in context. A part that rarely fails can still be highly critical if replacement takes weeks.
The Criticality Matrix: Keep It Simple
Rather than turning spare planning into a sprawling inventory debate, the guests recommend a simple criticality matrix based on two variables: consequence and time to recover. That framework keeps decision-making practical.
If a part has high consequence and a long replacement timeline, it deserves immediate attention. If it has lower consequence or can be obtained quickly, the urgency changes. The value of the matrix is that it moves the discussion away from opinion and toward a common operating framework that both operations and finance can understand.
That matters because spare-parts planning often stalls when one side sees only cost while the other sees only risk. A one-page matrix helps both sides align faster. It does not have to be elegant. It only has to help the team answer the question that matters most during a failure: what breaks if this goes down, and how fast can we recover?
Stock or Source? The Real Trade-Off
The episode spends time on one of the most practical tensions in operations: whether to own spares or rely on fast procurement. James pushes back on blind vendor dependency, especially when teams assume a supplier SLA will hold under real-world conditions like holidays, factory shutdowns, or shipping disruptions. Michael agrees that stocking everything is not realistic. Owning spares ties up capital, consumes storage space, and creates obsolescence risk.
The rule they land on is useful because it is balanced. Own spares for true single points of failure where vendor lead times exceed your outage tolerance. For lower-risk items, rely on procurement only if those vendor commitments have actually been tested.
That last point is important. Quoted lead times are not the same as proven recovery capability. Michael recommends running a mock order, timing it, confirming delivery windows, and documenting both escalation paths and fallback sources. In other words, if your continuity plan depends on a vendor, that dependency should be validated before the outage happens, not during it.
Why Stored Spares Fail When You Need Them Most
A major theme in the discussion is that buying a spare does not finish the job. Stored parts can silently degrade into false confidence. Michael shares a case where a “new” controller from inventory did not match the firmware on the live network, wasting critical hours. That example shows why compatibility tracking matters just as much as the physical part itself.
The guests argue that spares should be treated as managed assets with lifecycle tags, not as forgotten emergency backups. Firmware versions, board revisions, and compatibility notes all need to live in the spare record. If a spare moves beyond its useful compatibility window, the team should update it, repurpose it, or retire it.
James brings in another overlooked factor: storage conditions. Humidity, temperature, batteries, and packaging affect whether a spare will actually work. Anti-static protection, dry storage, and clear labels are not administrative details. They are part of reliability. A spare sitting in the wrong environment may look available on paper while being unusable in practice.
Make Spares Part of CMMS, Not Tribal Knowledge
One of the strongest operational recommendations in the episode is to make spares first-class records in the CMMS. James lays out a simple structure: each spare should have an ID, a storage location, a last test date, a firmware version, and a scheduled test interval.
This is a meaningful shift. In many environments, spare information lives in a drawer, a spreadsheet, or one technician’s memory. That creates delay at exactly the moment a team needs speed. Putting spares into the CMMS makes them visible, trackable, and actionable.
James also recommends rotating spares into low-risk circuits during preventive maintenance. That keeps parts exercised and gives technicians real installation practice. The result is better than a theoretical emergency plan. It creates operational muscle memory. When something fails at 2:00 or 3:00 a.m., the team is not handling the spare for the first time.
Cross-training supports the same goal. If only one technician knows where a spare is stored or how to install it, the organization still has a single point of failure. Teaching at least one colleague how to find and replace the part is, as the episode puts it, cheap insurance.
Two Cases That Capture the Difference
The conversation includes two short examples that make the broader point memorable. In the first, a campus kept a spare network uplink module in a locked cabinet. When a contractor accidentally fried a live port at 2:00 a.m., the team hot-swapped the module, triggered failover, and restored connectivity in 20 minutes. No alarms escalated, and no tenant disruption followed. In that case, the spare paid for itself in a single event.
In the second example, a building had an aging BAS spare stored in a damp utility closet. When the primary controller failed, the spare would not boot because its connectors had corroded. Instead of saving the day, the spare extended the outage and contributed to days without proper HVAC control. The building eventually corrected the problem by improving storage, adding quarterly boot tests, and logging the spare in CMMS, but only after a painful failure exposed the gap.
Together, those examples define the difference between a real spare strategy and the illusion of one. Availability is not enough. Readiness is what counts.
Three Practical Moves to Make This Week
Before wrapping up, the guests leave listeners with three straightforward rules.
- Build a one-page criticality matrix for top systems so operations and finance can align quickly
- Put spares into the CMMS with IDs, test intervals, and rotation schedules
- Run a live vendor SLA test annually and document accountability plus fallback options
They also recommend a 15-minute audit this month on one system: UPS, network, BAS, or access control. Identify the vulnerable part, confirm who supplies it, and determine whether a tested spare exists. If not, add it to the CMMS or the next capital planning cycle and schedule a validation test.
Michael adds an especially practical finance tip: attach a downtime cost estimate to the CMMS record when justifying a spare. That makes future decisions easier because the conversation is grounded in business impact, not just equipment cost.
The Real Value of a Spare Parts Playbook
The biggest takeaway from this episode is that spare-parts planning is not a side task for maintenance teams. It is a business continuity discipline. The right spare, kept in the right condition and documented the right way, can turn a multi-day disruption into a brief service event. The wrong spare strategy, or no strategy at all, leaves teams reacting under pressure while tenants feel the impact.
For owners, operators, and facilities teams, this is a practical reminder that resilience often comes from unglamorous work done in advance: mapping criticality, testing assumptions, documenting inventory, and training people before the failure happens. If you want a better handle on where your building is exposed, this episode offers a smart place to start. Listen to the full conversation and use the framework to audit one critical system this month.