Show Notes
When Redundancy Looks Better on Paper Than It Works in Reality
A mid-rise office building experiences a power flicker. The generator starts, the facilities team initially feels relief, and then the building goes dark 15 minutes later. The backup generator is still running, yet the building has lost power.
That scenario captures the central issue in this episode: backup equipment alone does not create resilience. A generator, dual utility feeds, redundant network paths, or multiple chillers can all create confidence without guaranteeing continuity. The real question is whether every supporting component, handoff, procedure, and person will work when the primary system fails.
The discussion focuses on the difference between owning backup equipment and maintaining a building that can keep operating through an outage.
The Hidden Single Points of Failure
The most common redundancy failure is a single point of failure concealed inside an otherwise redundant design.
- Two generators may share one fuel-transfer pump.
- Redundant chillers may depend on one condenser-water pump.
- Multiple internet connections may enter the building through the same conduit.
- A generator may run normally while a transfer switch fails to move building loads to backup power.
- Fuel may be available on site, but a clogged fuel line can keep generators from receiving it.
One example involved two diesel generators that were serviced and tested monthly. During a utility outage, Generator A ran for 15 minutes before shutting down. Generator B then ran for 10 minutes before shutting down. The building had 10,000 gallons of diesel fuel, but both engines depended on a single failed fuel-transfer pump. The fuel supply existed, the generators existed, and the maintenance records existed, but one shared component defeated the entire backup strategy.
The lesson is straightforward: redundancy must include the systems that support the backup equipment, not only the equipment itself.
Use Failure Mode Analysis to Find the Holes
Failure mode analysis is presented as a practical way to uncover weak points before an incident exposes them. The process requires asking uncomfortable questions:
- What happens if a valve sticks?
- What happens if a breaker trips?
- What happens if a shared pump fails?
- What happens if the transfer switch does not transfer?
- What happens if a backup service provider cannot be reached at 2 a.m.?
No building can analyze every possible failure at the same depth. The right level of effort should reflect the actual cost of downtime. A small office building does not necessarily need hospital-grade redundancy. But building size alone is not a reliable measure of risk. A small office may house a server room that supports a medical billing system, where downtime can create losses of thousands of dollars per hour.
The goal is to make risk-based decisions with a clear understanding of what interruption would cost the organization and its tenants.
Why Testing Must Include Real Failover
Starting a generator is not the same thing as proving that a building can survive a utility failure. Testing needs to confirm that critical loads actually move to backup power and remain supported under load.
- Review whether load testing is being performed.
- Confirm whether testing includes a real failover, not just generator startup.
- Validate the transfer switch and the full power path to critical systems.
- Test backup systems under conditions that resemble an actual outage.
- Conduct a real drill rather than relying only on a tabletop exercise.
A load test may cost a few thousand dollars, but it can expose a weakness that could otherwise produce a much larger operational loss. The inconvenience of scheduling and performing the test is minor compared with discovering a failed transfer switch, clogged fuel path, or unavailable service contact during an unplanned outage.
A Practical Backup-System Audit
Property and facilities leaders can begin with a disciplined audit of their critical systems.
- Map critical loads, including life safety, elevators, and IT systems.
- Trace the power path from the utility source to every critical load.
- Identify every shared component and potential single point of failure.
- Map network paths as carefully as power paths.
- Review testing records and determine what has truly been validated.
- Document responsibility for each system and vendor relationship.
- Collect pictures of nameplates, transfer switches, and breakers.
- Verify that emergency contact numbers lead to people who will actually respond.
Documentation matters during an outage. Accurate records of equipment, ownership, contacts, and system locations help teams act quickly rather than searching for basic information in a high-pressure situation.
Key Takeaway
Redundancy is not a product purchase. It is an operating capability built through testing, documentation, coordination, and ongoing maintenance. Before the next storm or utility disruption, trace the hidden dependencies in your building, run a meaningful load test, and make sure your backup plan is prepared for the real world—not just the diagram.
The Redundancy Illusion in Building Operations
Backup generators, dual power feeds, redundant network paths, and secondary cooling systems are intended to protect buildings from disruption. They also create a risk that is easy to overlook: confidence in systems that have never been tested as a complete chain.
A building can appear highly resilient on paper and still fail during a routine outage. Consider a mid-rise office building where power flickers, the generator starts, and the facilities team assumes the problem is contained. Fifteen minutes later, the building goes dark. The generator continues running, but the building is not receiving the power it needs.
How can that happen? The generator may be healthy while the transfer switch fails to transfer the load. Fuel may be present but unable to reach the engine because of a clogged line. The problem may not be the backup asset everyone talks about. It may be a supporting component no one identified, exercised, or documented.
That is the redundancy illusion: treating the presence of backup equipment as proof of resilience.
Backup Assets Are Only One Part of the System
Redundancy is often discussed in terms of visible equipment. Two generators sound safer than one. Dual internet circuits sound safer than a single carrier. Redundant chillers sound safer than relying on one unit.
But critical systems do not operate as isolated assets. They depend on fuel delivery, pumps, switches, breakers, conduits, cooling, network paths, vendors, procedures, and people. If any shared dependency fails, the redundancy may disappear at the exact moment it is needed.
One example illustrates the problem clearly. A building had two diesel generators. Both were serviced monthly and tested monthly. When the utility supply dropped, the first generator started and ran for 15 minutes before shutting down. The second generator started and ran for 10 minutes before shutting down as well.
The building had 10,000 gallons of diesel available. The issue was not a lack of fuel, nor was it a failure of either generator. Both generators depended on one fuel-transfer pump, and that pump had failed. A single shared component starved both engines.
The same logic applies to cooling. A property may have redundant chillers but only one condenser-water pump. If that pump fails, both chillers can trip. The building has multiple major assets, yet a single supporting device can disable them all.
Network redundancy deserves the same scrutiny. Two internet connections do not provide meaningful path diversity if both enter the building through the same conduit. A damaged conduit can take down both connections at once, regardless of how many service providers appear on the invoice.
Find Shared Dependencies Before an Outage Does
Failure mode analysis provides a disciplined way to expose the hidden choke points in a building’s backup design. It does not require predicting every imaginable problem. It requires asking targeted questions about what could interrupt the systems that must remain operational.
Start with questions such as: What happens if this valve sticks? What happens if this breaker trips? What happens if this pump fails? What happens if the transfer switch does not move the load? What happens if the backup vendor cannot be reached after hours?
These questions can be uncomfortable because they reveal that a system may be less resilient than assumed. That discomfort is valuable. It is much easier to address a weakness in a scheduled maintenance window than during an outage affecting tenants, life-safety functions, elevators, IT operations, or revenue-producing systems.
The purpose is not to demand the same infrastructure from every building. A small office building does not automatically need the level of redundancy expected in a hospital. The appropriate investment should reflect the consequence of downtime.
However, leaders should not use building size as a shortcut for risk. A small office may host a server room supporting a medical billing system. If that system becomes unavailable, the organization may lose thousands of dollars per hour. The operational impact—not the square footage—should shape the resilience conversation.
Testing Is Not the Same as Starting Equipment
Many backup programs fail because their tests validate only a narrow portion of the design. Starting a generator proves that the generator can start. It does not prove that the transfer switch will operate, that critical loads will transfer, that fuel will continue flowing, or that the building can sustain operations under load.
A meaningful test validates the actual failover path. It confirms that the building can move from the utility source to backup power and continue supporting the systems identified as critical. It also reveals whether support teams, service providers, and internal contacts can respond when the test uncovers an issue.
Load testing can be inconvenient and carries a cost. But the comparison should not be between the test cost and zero cost. The proper comparison is between the cost of controlled validation and the potential cost of uncontrolled downtime.
A few thousand dollars spent on a load test can reveal a single point of failure that could otherwise result in a much larger outage. The risk may involve lost business, tenant disruption, IT downtime, delayed billing, safety concerns, or expensive emergency response. Testing is not a paperwork exercise; it is an operational investment.
A Practical Audit for Property and Facilities Teams
A strong resilience review begins by mapping what truly has to stay online. This should include life safety, elevators, IT systems, and any other operationally critical loads.
Then trace the full power path from the utility source to each critical load. Do not stop at the generator. Include transfer switches, breakers, fuel systems, pumps, control components, and every other supporting dependency. The same process should be applied to network connectivity. Trace each provider’s path into the building and determine whether supposedly diverse circuits share a conduit or another common failure point.
Next, review the testing record. Determine whether the organization has conducted load tests and whether failover has been tested in reality. A tabletop discussion can improve communication, but it is not a substitute for throwing the switch and observing what happens when the utility feed is unavailable.
Ownership is equally important. During a 2 a.m. outage, who is responsible for calling the electrician, generator service company, or fuel supplier? Are current contact numbers available? Does the backup number reach a real person, or does it route to voicemail because the primary technician is unavailable?
Documentation supports fast decision-making under pressure. Keep records of system ownership, emergency contacts, nameplates, transfer switches, and breakers. Take pictures while conditions are normal. During an outage, those details can save valuable time and reduce guesswork.
Resilience Is a Capability You Maintain
True redundancy is not something a building purchases once and forgets. It is a capability maintained through testing, documentation, coordination, and clear accountability.
A building may own backup equipment and still be unprepared. A resilient building knows its critical loads, understands its single points of failure, has tested its failover path under meaningful conditions, and can reach the right people when something breaks.
Before the next storm or utility disruption, conduct a load test, trace shared dependencies, and have the uncomfortable conversation about what could fail. The objective is not to accumulate more backup equipment. It is to keep the building working when everything else falls apart.
Listen to this episode of Built, Wired & Secured for a practical discussion of hidden failure points, testing priorities, and the operational discipline required to make redundancy real.