GDS Technology — Built, Wired and Secured podcast banner
Watch on YouTube →
Episodes Built
Episode 53

Proof in the Power: Validating Building Resilience Through Practical Tests

June 23, 2026
Key takeaways
  • Resilience is proven through observable, timestamped, end-to-end evidence rather than individual device status or technician notes.
  • Hidden dependencies, stale documentation, legacy circuits, static routes, and physical patching can cause critical services to fail even when power or routing appears healthy.
  • Tabletop exercises support communication and dependency mapping, while targeted live tests are required to validate physical behavior.
  • Generator stabilization time, UPS runtime, load segregation, and transfer sequencing must be tested together against business impact requirements.
  • Test findings should be treated as defects with a single owner, a documented mitigation or budget decision, and tracked follow-through.

Show Notes

When the Generator Starts but the Building Still Fails

A building can appear prepared for an outage and still leave tenants in the dark. This episode opens with a practical failure: backup power was activated, yet emergency lights stayed off, elevators stalled, card readers failed, and the leasing desk was overwhelmed. The generator itself had started. The issue was that a transfer path had changed and emergency-lighting circuits were no longer included in the protected scope.

The lesson is simple: resilience is not proven when an individual device reports healthy. It is proven when the complete path works from the source of power or connectivity through the systems people actually use.

What “Proving Resilience” Means

Resilience should be observable, timestamped, and tested end to end. A note that says a breaker moved or a generator started is not enough. Teams need evidence that the intended source-to-load path operated successfully within an acceptable time.

  • Capture logs, dashboard screenshots, photos, and timestamps.
  • Use signed checklists and identify who was responsible for each test step.
  • Validate transfer switches, UPS behavior, carrier routing, physical patching, and critical services together.
  • Test the systems tenants depend on, not only the devices technicians can see in a console.

These artifacts turn a test from an anecdote into an operational record. They also make handoffs stronger, help justify remediation work, and give leadership better information for capital planning.

Why Building Systems Drift Out of Sync

Many resilience problems are caused by hidden dependencies and stale assumptions rather than one dramatic technical failure. Documentation may describe a protected circuit, a failover route, or a patching arrangement that no longer matches the building. Renovations, panel reroutes, changed maintenance schedules, legacy circuits, and incomplete runbook updates can create a gap between the map and the real environment.

One example in the episode involved a mixed-use campus where primary fiber was removed to validate failover. Routing converged successfully, so the network looked healthy from a BGP perspective. But phones remained down and the security center lost camera feeds. Certain services were still tied to a legacy circuit that did not fail over because of static routes and a mispatched panel.

That test did not reveal a routing problem. It revealed an end-to-end ownership and governance problem. If no one owns the outcome across facilities, IT, vendors, carrier services, and building systems, a test can become a checkbox exercise that creates false confidence.

Tabletops and Live Tests Have Different Jobs

Tabletop exercises are valuable. They help teams review communications, clarify roles, identify likely dependencies, plan escalation paths, and prepare a safe test. They do not, however, prove that physical infrastructure and connected systems will behave as expected.

Live tests provide that physical validation. They can be limited in scope and scheduled carefully to reduce disruption:

  • Choose a low-impact maintenance window.
  • Test a single critical path instead of attempting an uncontrolled building-wide event.
  • Pre-stage rollback plans before changing system state.
  • Notify tenants and coordinate needed safety permits.
  • Confirm facilities, IT, and applicable vendors are available for the test.

The recommended approach is not tabletop versus live testing. Use a tabletop to plan and coordinate, then use a targeted live test to validate the real behavior. Small, frequent tests are more useful than rare, oversized exercises that are difficult to schedule and act on.

Power Resilience Requires Sequencing, Not Assumptions

Generator and UPS capacity must be evaluated together. In one example, a UPS was sized for immediate IT loads while the generator had been sized and tested under light conditions. During a real outage scenario, the generator took longer than expected to stabilize. UPS batteries began dropping sooner than planned, and non-IT equipment—including strobe lights and remote-access panels—lost power.

There was no universal answer. The organization could increase UPS runtime, move non-critical loads off the UPS, or retime generator load transfer. The right choice depends on business impact: what fails, who is affected, and how long the organization can tolerate degradation. Testing creates the facts needed to decide between capital investment and an operational workaround.

A Practical Resilience Test Checklist

  1. Confirm ownership. Obtain facilities and IT signoff, and involve the vendors responsible for the path being tested.
  2. Select one critical path. Keep the scope tight and define a one-sentence success criterion with a time limit.
  3. Capture evidence. Record timestamps, logs, photos, and a responsible-party signature.
  4. Validate rollback and debrief. Confirm the return path works, then hold a 15-minute debrief with actions and next steps.

Repeat this process quarterly for critical paths. A concise test with clear acceptance criteria is more actionable than a broad test that produces vague findings.

Turn Test Results Into Action

Testing only improves resilience when a discovered weakness is treated as a defect. Assign a single owner. Decide whether the response is a funded capital improvement, an operational mitigation, a vendor-SLA item, or an update to protection scope and sequencing. Track it until it is closed.

For a practical starting point, take three actions next week: run one targeted live test during a low-impact window, hold a short tabletop with facilities, IT, and primary vendors, and create or update a one-page checklist with acceptance criteria and an assigned owner. The goal is to replace hope with evidence—and make the results useful before a real incident forces the issue.

Deeper dive

Building Resilience Is Not a Document: It Is a Demonstrated Outcome

Many commercial properties have outage plans, generator maintenance records, carrier contracts, and emergency procedures. Those are important foundations. But they do not necessarily prove that the building will continue operating when power or connectivity is interrupted.

A resilience program fails when it confirms isolated components rather than validating the complete experience of the people who depend on the building. A generator can start while emergency lighting remains dark. Network routing can converge while phones, card readers, or camera feeds are unavailable. A UPS can be healthy during a light test but still run out of usable runtime while a generator stabilizes under actual conditions.

The difference between having a plan and having resilience is evidence. Building owners, facilities leaders, and IT managers need practical ways to prove that critical systems work together, within the time their business and tenants can tolerate.

Start With the Path, Not the Device

It is easy to test a generator, a transfer switch, a UPS, or a carrier circuit as separate assets. It is harder—and more valuable—to test the path from a source through the systems that serve occupants and operations.

For backup power, that path may begin at the generator and transfer equipment, continue through electrical panels and UPS equipment, and end at emergency lighting, access control, elevators, remote-access panels, or essential IT systems. For connectivity, it may begin with a primary carrier failure, move through routing and failover equipment, and end with phones, cameras, security operations, and tenant-facing services.

A component test can still be useful, but it cannot stand in for end-to-end validation. If a technician records that a breaker moved, that record says little about whether a particular protected load stayed online. If a network dashboard says failover routing converged, it does not establish that a legacy circuit, static route, or physical patch panel did not interrupt a service further downstream.

Define the test around a real dependency. Ask: what service matters, what supports it, and what must happen for the service to remain available or return within an acceptable period?

Hidden Dependencies Are the Most Common Surprise

Building environments change continuously. Renovations alter panel arrangements. Legacy circuits stay in service longer than expected. Equipment gets moved, a patch is changed, a static route remains in place, or a maintenance schedule changes. Documentation may still reflect the previous environment.

That is how the map and the territory diverge. The risk is not necessarily poor maintenance; it is that multiple teams and vendors can make sensible local changes without anyone validating the full system outcome afterward.

Consider a campus failover test where primary fiber is intentionally removed. BGP convergence may show that the network successfully selected another path. Yet phones and security-center camera feeds may remain unavailable because some services are hardwired to a legacy circuit that never switched. In that situation, the routing view is accurate but incomplete. The failure is found in the physical patching, static routing, and ownership boundaries that were not included in the original assumption.

That is why resilience needs governance as well as technology. Someone must own the outcome across facilities, IT, carriers, security systems, and vendors. Without that ownership, an otherwise well-intended test becomes a checkbox event that records activity without proving service continuity.

Use Tabletop Exercises to Prepare; Use Live Tests to Prove

Tabletop exercises have a meaningful role in resilience planning. They help teams identify stakeholders, practice communication, map expected dependencies, clarify escalation paths, and discuss rollback decisions. A tabletop can expose confusion before a maintenance window begins.

But it cannot confirm physical behavior. It cannot prove that a transfer path includes the expected circuit, that a UPS carries the intended load long enough, or that a carrier failover restores every operational service. Only a live test can do that.

Live testing does not have to mean a broad, high-risk event. The practical approach is to make it small, intentional, and controlled:

  • Schedule during a low-impact window.
  • Choose one critical end-to-end path.
  • Set clear acceptance criteria before beginning.
  • Notify affected tenants and coordinate appropriate safety permits.
  • Pre-stage the rollback plan and identify who can authorize it.
  • Ensure facilities, IT, and key vendors are available.

This approach reduces tenant impact while still producing evidence about how the actual environment behaves. It is more useful to test one meaningful path every quarter than to wait for a large annual exercise that is difficult to coordinate and too broad to diagnose effectively.

Power Resilience Depends on Timing and Load Decisions

Generator and UPS planning often creates false confidence because each system may appear correctly sized in isolation. The critical question is how they perform together during the transition.

Imagine a site where the UPS supports immediate IT loads and the generator has been tested under light conditions. During an outage, the generator may take longer than expected to stabilize. The UPS batteries may begin to dip before the transfer sequence is complete. Systems that were assumed to be protected—such as strobe lights or remote-access panels—can drop unexpectedly.

There are several possible remedies. The organization might increase UPS runtime. It might move non-critical loads away from the UPS. It could retime generator load transfer or alter the protection scope. Each option has a different cost and operational consequence.

The decision should begin with business impact. What fails if the system is unavailable? Which tenants or operations are affected? How long can the service be degraded? The answers determine whether a capital improvement is justified or whether an operational workaround is acceptable. Testing replaces assumptions with data that can support that decision.

Make Acceptance Criteria Specific and Observable

A useful resilience test begins with a short success statement that includes a time limit. For example, a team might define that a particular critical service must remain available or return within a stated period after a controlled failover. The episode’s guidance is to keep this to one sentence and make it observable.

Then capture the proof. Evidence should include timestamps, logs, dashboard captures, photos where relevant, and a signature from the responsible party. These records are not bureaucracy. They establish what actually happened, reduce ambiguity during future handoffs, and make it easier to explain remediation needs to leadership, vendors, or capital-planning stakeholders.

The evidence should describe the whole story: the planned condition, the observed result, any deviation from the acceptance criterion, rollback status, and the next action. This is particularly important when a test appears successful at one layer but reveals a failure elsewhere.

Do Not File Failures Away—Treat Them as Defects

The most damaging version of resilience testing is not a test that finds a weakness. It is a test that finds a weakness and then leaves the result in a shared drive with no owner or decision.

When a test exposes a gap, triage it as a defect. Assign one person to close the loop. Determine whether the response is a capital request, a configuration or sequencing change, an operational mitigation, a vendor-SLA issue, or a documentation update. Then track it until it is resolved.

This discipline turns testing into a business process rather than an annual compliance artifact. It gives property teams a clearer basis for prioritizing improvements, protects tenants from avoidable surprises, and helps leaders direct spending toward issues that have been demonstrated rather than merely suspected.

Three Actions to Take Next Week

  1. Run one targeted live test. Choose a critical path, use a low-impact window, and capture timestamps and logs.
  2. Hold a short tabletop. Bring facilities, IT, and primary vendors together to identify likely hidden dependencies before the live test.
  3. Create a one-page checklist. Include the acceptance criterion, evidence requirements, rollback steps, and a named owner for follow-up.

These are modest actions, but they establish the operating rhythm that resilience requires. The aim is not a perfect, oversized exercise. It is a repeatable process that finds gaps early, produces useful evidence, and ensures someone acts on the result.

For a deeper discussion of targeted live tests, UPS and generator sequencing, carrier failover, hidden dependencies, and evidence-based follow-up, listen to this episode of Built, Wired & Secured. The core message is practical: small, frequent, evidence-based tests are stronger than rare exercises that create confidence without proof.