GDS Technology — Built, Wired and Secured podcast banner
Watch on YouTube →
When Parts Disappear: Managing Spares, Firmware, and End‑of‑Life in Building Systems
Episodes General
Episode 75

When Parts Disappear: Managing Spares, Firmware, and End‑of‑Life in Building Systems

July 11, 2026
Key takeaways
  • A missing spare part is only part of the risk; firmware, configuration, and ownership gaps make recovery slower and harder.
  • Early warning signs include outdated spare lists, undocumented firmware, vendor-only notes, and procurement driven only by lowest cost.
  • Vendor managed spares can work for lower impact items, but high impact or end-of-life components often need on site spares and firmware custody.
  • Lightweight bench testing and attaching acceptance results to spare records can prevent major surprises during outages.
  • A 90-day review should identify top impact assets, vault firmware and configs, map obsolescence, and assign accountability.

Show Notes

Why small failures create big building problems

This episode of Built, Wired, and Secured starts with a story that makes the whole issue real fast: a rooftop AHU dropped offline in the middle of the night because a controller EROM failed, and the replacement that had already been ordered was discontinued. Within hours, tenants were complaining about heat. The team patched it temporarily, but the bigger lesson was clear. In building systems, outages do not always begin with a major disaster. Sometimes they begin with one missing part, one missing firmware image, or one undocumented dependency.

Alex Morgan frames the conversation around the quiet risks that sit in the background until they become urgent: disappearing spares, orphaned firmware, outdated inventories, and procurement habits that focus only on lowest cost. Michael Harrington and James Rogers explain why these issues create real operational and business risk for facilities teams, IT leaders, and property owners.

What makes a spare part a single point of failure

A major theme in the episode is that the risk is almost never the part by itself. The real risk is the combination of the hardware, its firmware, its configuration, and the knowledge required to restore it. If any one of those elements is missing, a simple replacement can turn into a prolonged outage.

  • A failed controller can quickly become a tenant comfort issue.
  • Comfort complaints can turn into a work order backlog.
  • That backlog can create lease friction and operational strain.
  • Expedited shipping, overtime, and temporary workarounds can cost far more than the part itself.

The conversation makes a practical point that many owners overlook: a cheap component can still sit at the center of a very expensive recovery event.

Early warning signs teams should not ignore

The episode lays out several red flags that suggest a building or facility is more exposed than it looks on paper. These are not theoretical issues. They are operational clues that a team may be one failure away from a long night.

  • Maintenance notes that say “vendor only” with no internal custody plan.
  • Spare parts lists that have not been updated in years.
  • Controllers running older firmware with no recorded images.
  • No clear answer to who holds the last known good firmware.
  • Procurement decisions driven only by the cheapest SKU or a single reseller.

That last point matters. The group argues that procurement is often treated as cost control when it should also be treated as risk management. A lower purchase price can increase fragility if it narrows sourcing options or leaves a site exposed when a product line reaches end of life.

Vendor managed spares versus on site spares

One of the more useful parts of the discussion is the tension between vendor managed inventory and on site spare strategy. The answer is not either-or. It depends on risk.

Vendor managed spares can reduce carrying cost for commodity parts while a product is still supported. But the episode makes it clear that this is not a complete end-of-life strategy. Vendor stock does not guarantee long-term availability, and contracts rarely force a vendor to provide firmware custody or replacements years down the road.

The practical recommendation is to match the approach to impact:

  • Use vendor managed spares for lower impact production items.
  • Keep on site spares for high impact, single source, or discontinued items.
  • Pair critical spares with firmware custody and documented configuration history.

Three low cost moves teams can start now

Rather than turning this into a multi year transformation project, the episode stays focused on things teams can do in weeks.

  • Inventory the truly critical items, especially controllers, gateway modules, and sensors that can trigger tenant impact cascades.
  • Vault firmware and configurations with clear version labels and device IDs.
  • Build a rolling obsolescence timeline so procurement knows when last-time-to-buy decisions need to happen.

The group also emphasizes that validation does not require a full lab. A lightweight bench setup can be enough. Their example is simple: a powered bench, a small switch, and a script that verifies firmware boots and basic I/O. If the acceptance test result is attached to the spare record, the team avoids guessing during an outage.

Why ownership matters under pressure

Another practical recommendation is to assign responsibility in advance. When a system fails, confusion about ownership wastes time. The speakers recommend a simple table that answers who owns spares, who owns firmware custody, and who owns acceptance testing. Depending on the environment, that could be facilities, IT, or a vendor. What matters is that the answer exists before the emergency starts.

They also stress the value of documenting the rationale behind decisions. If a team chooses not to hold an on site spare, that choice should still be written down with the accepted risk, recovery plan, and vendor lead times. That turns a vague assumption into a defensible decision.

A real example of time saved

The payoff comes through in an anonymized example from the field. After a storm, an access control headend failed and the listed replacement was discontinued. The team sourced a board from another facility, backed up its firmware, and restored access in under eight hours. They compare that to a prior similar event that stretched close to forty eight hours. The savings were not abstract. They showed up in tenant hours preserved and overtime avoided.

They also describe a quarterly routine of pulling configs and firmware and bench testing one spare. It takes about a day each quarter and has prevented several emergency procurements.

The 90 day review to run this week

The episode closes with a prioritized review that listeners can begin immediately. The message is straightforward: do not wait for the next outage to learn what your building depends on.

  • Identify the top impact assets first.
  • Focus spare and firmware effort on those assets.
  • Vault firmware and configuration with versioning.
  • Attach acceptance test results to spare records.
  • Create an obsolescence timeline and a last-time-to-buy trigger.
  • Assign accountability for spares, firmware custody, and the test bench.
  • Document recovery plans so future decisions are objective.

The final takeaway is simple and useful. Ask early: what breaks if this goes down? Then plan around the answer. The team argues that a little preventive work up front keeps nights quieter later, and this episode makes a strong case that they are right.

Deeper dive

When small building system failures turn into major operational problems

Most outages do not start with something dramatic. They start with something easy to overlook.

In this episode of Built, Wired, and Secured, Alex Morgan opens with a story that captures the problem perfectly: a rooftop AHU dropped offline in the middle of the night because the controller’s EROM failed, and the replacement that had already been ordered was discontinued. Tenants were complaining about heat within hours. The team found a temporary patch, but the clock was already running.

That example sets the tone for a practical conversation with Michael Harrington and James Rogers about disappearing spare parts, firmware drift, and end-of-life risk in building systems. The point is not that parts fail. Everyone knows they do. The point is that many organizations still treat spare inventory, firmware custody, and lifecycle planning as minor administrative tasks when they are really core resilience issues.

If you own, operate, or support buildings, this discussion offers a clear way to think about why small technical gaps create outsized operational and financial consequences.

The problem is rarely the part by itself

One of the strongest ideas in the episode is that a failed part is almost never the full story.

When a controller, module, or sensor goes down, the recovery depends on more than hardware availability. It also depends on the configuration tied to that device, the firmware version it needs, the ability to validate a replacement, and the institutional knowledge required to restore service. Lose any one of those, and what should have been a straightforward replacement can turn into an extended outage.

That is why the speakers describe these events as “quiet failures.” The risk is present long before the incident. It builds slowly through neglected inventories, incomplete records, firmware images that were never preserved, and assumptions that the vendor will always be able to help later.

Then the outage happens, and a low cost component suddenly becomes the center of a high cost response.

Why tenants feel these failures immediately

The episode does a good job connecting technical details to business outcomes. A failed controller is not just a line item on a maintenance report. In the rooftop AHU example, tenants felt the impact right away because comfort was affected. That turned a device issue into a service issue.

As Michael and James explain, the sequence usually looks something like this:

  • A building system component fails.
  • Occupants notice the effect quickly.
  • Complaints and work orders begin to stack up.
  • Temporary fixes consume staff time.
  • Expedited shipping, overtime, and workaround costs rise.
  • The issue starts affecting tenant confidence and potentially lease relationships.

The hardware may be relatively cheap. The business impact usually is not.

That framing matters because it changes how owners and operators should evaluate spare parts and end-of-life planning. This is not only about technical housekeeping. It is about protecting occupancy, responsiveness, and continuity.

The early warning signs most teams already have

The conversation stays grounded in the kinds of signals teams can identify without a major consulting engagement. Several red flags come up that should prompt immediate review.

  • Maintenance notes that say “vendor only.”
  • Spare lists that have not been touched in years.
  • Controllers running older firmware with no recorded images.
  • No clear answer to who holds the last known good firmware.
  • Buying decisions based only on the cheapest SKU or a single reseller.

These are simple indicators, but they reveal a lot. If a team cannot quickly identify where the approved firmware lives, whether a critical spare has been validated, or how replacement lead times change near end of life, the environment is more fragile than it appears.

The procurement point is especially important. The episode argues that organizations often make purchasing decisions as if price is the only variable that matters. In practice, source concentration and support horizon matter too. Saving money on the front end can create a more brittle operation on the back end.

Vendor managed spares are useful, but not enough

There is a useful debate in the episode about vendor managed spares. One side sees clear value in vendor managed stock for commodity items because it reduces carrying cost. The pushback is that vendor managed inventory does not solve end-of-life exposure.

That distinction matters. A vendor can help while a product line is active and supported. Once the line is sunset, the existence of a contract does not necessarily guarantee replacement availability, firmware handoff, or long-term support. As the speakers put it, contracts help, but they rarely force a vendor to preserve everything you need for years.

The practical answer is not to reject vendor managed programs. It is to use them where they fit and not confuse them with a full continuity plan.

  • Vendor managed spares make sense for lower impact items.
  • On site spares make more sense for high impact items.
  • Single source and discontinued components deserve extra attention.
  • Firmware custody should sit alongside the spare strategy, not apart from it.

That last point is one of the clearest takeaways from the episode. A spare without usable firmware or a known configuration path may not be much of a spare at all.

What teams can do in the next few weeks

The conversation avoids turning this into an abstract maturity model. Instead, it offers a short list of immediate moves that most teams can start without a large budget.

First, identify the truly critical assets. Not every device deserves the same level of attention. The focus should be on controllers, gateway modules, sensors, and other components whose failure creates tenant impact cascades or blocks recovery.

Second, vault firmware and configuration files with clear version labels and device IDs. The episode is direct here: if you cannot say who has the last known good firmware, you are exposed.

Third, build a rolling obsolescence timeline. Procurement should know when products are approaching end of life and when last-time-to-buy decisions need to happen. Waiting until the failure occurs is the most expensive time to have that conversation.

You do not need a full lab to validate spares

One practical detail that stands out is the discussion around testing. Many teams hear “validation” and assume they need a full lab environment. The speakers push back on that. Their view is that lightweight bench testing is enough to reduce risk meaningfully.

The example is simple: a powered bench, a small switch, and a script that confirms firmware boots and basic I/O functions. Then attach the acceptance result to the spares record.

That is a small operational habit, but it changes the quality of response during an incident. Instead of wondering whether a spare works, the team already has a record showing that it was tested and when.

For organizations managing multiple sites, that kind of habit scales better than ad hoc heroics.

Why ownership should be assigned before the emergency

Another strong recommendation from the episode is to define ownership clearly. Who owns spares? Who owns firmware custody? Who owns acceptance testing? Is that facilities, IT, or a vendor?

Under normal conditions, those questions feel administrative. Under outage conditions, they become urgent. If no one owns the answer, teams waste time, duplicate work, and start pointing fingers.

The speakers recommend a simple responsibility table and, just as importantly, written rationale behind each decision. If a team decides not to keep an on site spare, that should still be documented with the accepted risk, the recovery plan, and vendor lead times. That way the choice remains objective and defensible instead of becoming an emotional argument during a failure.

The operational payoff is real

The episode includes an anonymized example that shows why these habits matter. After a storm, an access control headend failed, and the listed replacement was discontinued. The team sourced a board from another facility, backed up its firmware, and restored access in under eight hours. They compare that result with a prior similar event that lasted close to forty eight hours.

The difference was not luck. It came from preparation and recovery discipline.

They also mention a quarterly routine that takes about a day each quarter: pull configs and firmware, then bench test one spare. That routine has already prevented several emergency procurements. For most organizations, that is a strong trade: one day of planned effort versus repeated high pressure purchasing and outage recovery.

A 90 day review that can pay off quickly

The closing recommendation is easy to act on. Start a 90 day review now.

  • Identify the top impact assets in your environment.
  • Confirm who has custody of firmware and configuration.
  • Vault images with versioning and device IDs.
  • Verify critical spares and attach test results.
  • Create an obsolescence timeline.
  • Set last-time-to-buy triggers.
  • Assign accountability for spares, firmware custody, and testing.
  • Document recovery plans and accepted risk where no spare is held.

The guiding question from the episode is the right one: what breaks if this goes down?

Once a team answers that honestly, the next steps become much clearer. Some assets justify on site spares. Some can rely on vendor managed inventory. Some need better documentation more than new hardware. The important thing is to move from assumption to policy.

For property owners, facilities leaders, and IT teams, this episode is a reminder that resilience often depends on unglamorous habits. Firmware custody, spare validation, lifecycle tracking, and ownership tables do not feel urgent until the moment they are.

By then, it is late.

If you want a practical place to start, listen to the episode and use its 90 day review framework to audit your own environment. A small amount of preventive work now can save a long outage, a rushed procurement, and a very uncomfortable conversation later. Listen here: https://builtwiredsecured.com/episodes/when-parts-disappear-managing-spares-firmware-and-end-of-life-in-building-systems