GDS Technology — Built, Wired and Secured podcast banner
Watch on YouTube →
Episodes Built
Episode 54

When the Cloud Goes Quiet: Managing Cloud Dependencies in Building Systems

June 24, 2026
Key takeaways
  • Cloud portals, cloud-mediated control, and cloud telemetry can create cascading building-system failure modes.
  • A system that works locally may still fail operationally if cloud-based credential sync or automation logic is unavailable.
  • Classify building systems as critical, important, or amenity to align local-control requirements with real operational impact.
  • Cloud-outage scenarios should be demonstrated during commissioning in the actual building environment before handover.
  • A 30-day cloud-touchpoint inventory and one blackout exercise are practical first steps for lean building teams.

Show Notes

When Cloud Convenience Becomes an Operational Risk

Modern commercial buildings increasingly depend on cloud services to manage access control, building automation, CCTV, elevator monitoring, telemetry, analytics, and vendor support. That centralization can provide better visibility, faster remote troubleshooting, and easier configuration changes. It can also introduce a hard operational question: what continues to work when the cloud service, identity platform, or vendor portal is unavailable?

In this episode of Built, Wired & Secured, Alex Morgan and Michael Harrington examine what happens when cloud-dependent building systems go quiet. The discussion centers on tenant continuity, safe access, local control, acceptance testing, vendor accountability, and practical ways owners and operators can improve resilience without over-engineering every system.

The First 30 Minutes of a Cloud Outage

When cloud-connected building systems fail, the first signs are often operational rather than technical. Badge readers may stop responding. Elevator status may disappear. The lobby may fill with people who cannot enter. Front desk radios may turn into a stream of access complaints.

Michael describes the immediate triage questions building teams should answer:

  • Can authorized people still get into the building?
  • Are life-safety systems still reporting as expected?
  • Do critical contractors have local access?
  • Is the issue a local network problem, a cloud-service outage, or a vendor outage?
  • Can teams maintain safe, predictable operations while the root cause is investigated?

The first half hour should focus on preserving access, maintaining safety, and giving the team a predictable operating path. That means understanding which building functions operate locally, which require the cloud, and where hidden dependencies may interrupt a process that otherwise appears resilient.

Three Common Cloud-Dependency Patterns

The episode identifies three patterns that frequently create fragility in building technology environments:

  • SaaS management portals: Vendor dashboards where rules, credentials, settings, and system administration live.
  • Cloud-mediated device control: Devices that connect to cloud services for commands, configuration, or operational changes.
  • Analytics and telemetry pipelines: Platforms that centralize monitoring data, analytics, or decision-making logic in the cloud.

Each pattern can be useful on its own. The risk increases when two or more are combined. A system may look operational because devices remain powered and connected locally, while a cloud-based synchronization process prevents new credentials, configuration changes, or automated actions from taking effect.

Hidden Sync Points Can Disrupt Tenant Access

A key example involves access control. Readers may be able to authenticate existing badges locally during a cloud outage, giving the appearance that the system has a solid offline mode. But if badge credentials are pushed from the cloud during tenant shift changes, a cloud outage can prevent new hires or newly authorized people from entering.

The failure is not necessarily at the reader. It is at the hidden synchronization point between the cloud and the access-control environment. That distinction matters because building teams need to test the actual workflows that matter: scheduled credential updates, contractor access, emergency access, local logging, and the restoration of normal operations.

Classify Systems by Operational Impact

Owners do not need to make every building system fully independent from the cloud. Instead, the episode recommends categorizing systems according to their impact on operations:

  • Critical: Life safety, core tenant access, and access required by emergency responders. These systems need local control and local logging.
  • Important: Systems that materially affect operations but can tolerate a defined interruption. These should have hybrid modes and clear failover expectations.
  • Amenity: Functions that can remain cloud-first because temporary loss is acceptable.

This classification should inform budget decisions, service-level expectations, commissioning requirements, acceptance tests, and vendor contracts. The central question is not whether cloud services are good or bad. It is whether a specific cloud function is mission-critical or a convenience that the building can lose for a defined period.

Make Fallback Practical, Not Theoretical

For owners concerned about cost or limited staff capacity, Michael recommends starting with one critical system. Rather than replacing an entire platform, implement a vendor-supported local fallback and require a measurable recovery expectation. One example used in the discussion is a target of less than 30 minutes to fail over to local control.

That incremental investment can be small compared with tenant disruption, emergency labor, extended lobby congestion, and the operational consequences of losing access or environmental control. The objective is not to eliminate all risk. It is to create a proven, usable fallback for the systems where downtime has the greatest consequences.

Test Cloud Failure Before Handover

Cloud-outage testing should be a scripted part of commissioning and acceptance, not an assumption documented in a sales presentation. A meaningful scenario includes:

  • Simulating a cloud outage.
  • Switching devices or controllers to local mode.
  • Confirming authorized users can still access the building.
  • Verifying that logs are written locally.
  • Measuring the time required to reach functional recovery.

If a vendor cannot demonstrate the fallback in the actual building network, owners can negotiate compensating controls such as redundant gateways, on-site controllers, or an emergency access method. If resilience has not been proven, handover may need to wait until it is.

Two Operational Lessons From Real-World Cases

The first case involved an office tower whose access control relied on a cloud identity service. During a morning tenant shift change, the identity service failed. Tenants could not badge in, the lobby became overwhelmed, and contractors could not reach critical floors. The team eventually used an undocumented local master-key mode, but nobody had practiced it. Afterward, the building added an on-site authenticated keypad fallback, documented the process, and tested it. Recovery time fell from hours to under 30 minutes.

The second case involved a campus HVAC analytics platform that stopped accepting telemetry for 48 hours. Automated cloud-cued setpoint changes did not execute, temperatures drifted, and occupants complained. The response was simple: local fallback schedules were added to controllers, and the team began conducting quarterly blackout exercises to validate schedule behavior. The building gained more predictable drift behavior and clearer escalation procedures.

A Five-Point Checklist for This Week

  • Inventory cloud touchpoints: List each system with a cloud dependency, classify it as critical, important, or amenity, and complete the inventory within 30 days.
  • Verify offline modes: Confirm what continues working locally, for how long, and whether the environment can fail over to local control in less than 30 minutes. Test at least a 24-hour window.
  • Add cloud failures to acceptance tests: Require cloud-outage demonstrations in the building environment before handover.
  • Run quarterly blackout exercises: Assign roles, establish communications, document outcomes, and improve the process after every exercise.
  • Maintain a living playbook: Document escalation contacts, local-log locations, fallback procedures, and tenant-notification templates.

The Minimum Viable Starting Point

For organizations with limited staff, the highest-return starting point is an inventory and one exercise. Complete the 30-day inventory, identify critical systems, and run a tabletop or simple blackout drill for one of them. That single exercise can expose whether the fallback is documented, whether it works, and where vendor support is required.

The Core Mindset Shift

Cloud capabilities can be valuable, but they should be treated as helpful additions unless they have been explicitly tested and guaranteed for operations. Design for autonomous operation first. Add cloud convenience on top. Make vendor demonstrations, documented fallback procedures, and routine exercises non-negotiable.

Deeper dive

When the Cloud Goes Quiet, Building Operations Cannot

Cloud-connected systems have become a normal part of commercial building operations. Access control platforms use cloud identity services and management portals. Building automation environments rely on remote dashboards, telemetry, and analytics. CCTV, elevator monitoring, and vendor support workflows increasingly depend on remote connectivity.

Those capabilities can make a property easier to manage. Centralized visibility can speed troubleshooting. Remote support can reduce delays. Cloud-based configuration can simplify administration across multiple systems or locations.

But a cloud-connected system is not automatically a resilient system.

When the cloud service behind a building function becomes unavailable, the consequences show up quickly in the physical world: tenants cannot badge into the building, contractors cannot reach critical floors, staff lose visibility into equipment status, and automated actions do not happen when they are expected to happen. The technology may be modern, but the operational question is straightforward: what still works when the cloud does not?

Cloud Dependency Is Often Hidden in the Workflow

The most important risk is not always that a device stops working outright. A building system may continue to operate locally while a less visible cloud-connected workflow fails.

Consider access control. Badge readers may authenticate existing users locally during a cloud outage. At first, that sounds like a strong fallback. But if a cloud service pushes new credentials during shift changes, newly hired employees or newly authorized contractors may still be locked out. The readers are online. The doors are functional. Yet a critical operational process has failed because the credential synchronization point was in the cloud.

This is why owners and operators need to look beyond a vendor statement that a platform has an “offline mode.” The better questions are:

  • Which functions continue working locally?
  • Which functions require a cloud service, identity platform, or vendor portal?
  • How long can the system operate in its local mode?
  • What workflows fail even if the devices remain operational?
  • How quickly can the building return to functional local control?

These questions reveal whether the cloud dependency is a convenience, a manageable operational limitation, or a true single point of failure.

Three Patterns That Create Fragility

Building technology environments commonly develop cloud dependency through three patterns.

First are SaaS management portals. These are vendor dashboards where administrators manage rules, credentials, configuration, and other system settings. They can simplify day-to-day administration, but they can also concentrate control in a service outside the building.

Second is cloud-mediated device control. In this model, devices connect to cloud services to receive commands or configuration changes. A controller may operate locally, yet changes to its behavior may depend on a remote service being available.

Third are cloud-based analytics and telemetry pipelines. These platforms centralize data and may also centralize operational logic. If an automated action is cued in the cloud, loss of telemetry or cloud processing can prevent that action from occurring.

Any of these patterns can be appropriate. The problem is not the presence of cloud technology. The problem is failing to understand what happens when multiple dependencies combine. A cloud management portal plus cloud credential synchronization can affect access. Cloud telemetry plus cloud-cued automation can affect HVAC performance. Centralized remote support plus a lack of local control can slow recovery when an outage occurs.

Start With an Operational Classification, Not a Technology Debate

Owners do not need a blanket policy that all systems must be locally managed. That approach can be expensive, difficult to maintain, and unnecessary for functions where temporary loss is acceptable.

A more practical approach is to classify systems by operational impact.

Critical systems include life safety, core tenant access, and access for emergency responders. These systems need local control and local logging because their availability directly affects safety, continuity, and emergency operations.

Important systems have a meaningful operational impact but may be able to tolerate an interruption for a defined window. These systems should use hybrid modes with documented failover expectations.

Amenity functions may be cloud-first when the property can tolerate their temporary unavailability. The point is to make this a deliberate decision rather than an accidental result of vendor architecture.

Once systems are classified, owners can connect resilience requirements to budgets, service expectations, acceptance tests, and contract language. Not every feature requires the same investment. The systems that control safe access and core operations deserve a higher standard.

A Practical Case for Local Fallback

Local fallback does not necessarily mean replacing an entire cloud platform or building a duplicate infrastructure. A practical first step is to select one critical system and implement a vendor-supported fallback that works in the real environment.

For access control, that may mean a local control path, an on-site authenticated keypad fallback, or an emergency access method that has been documented and tested. For HVAC operations, it may mean local schedules on controllers so that basic environmental control continues even when cloud analytics or cloud-cued setpoint changes are unavailable.

The financial case is operational. A lobby full of locked-out tenants, contractors unable to reach critical areas, extended emergency labor, occupant complaints, and unclear escalation paths all have costs. A limited, proven fallback can be far less costly than an outage that catches the building team unprepared.

One measurable expectation discussed in the episode is a failover to local control in less than 30 minutes. The exact target will depend on the system and property, but the important point is that the expectation should be measurable. “Resilient” is not a test. A documented recovery target is.

Commission for Failure, Not Just Normal Operation

Commissioning and acceptance testing should prove more than normal operation. They should prove that the building can function through a realistic cloud-service disruption.

A useful cloud-outage acceptance scenario includes disconnecting or simulating loss of the cloud service, switching devices to local mode, verifying access for authorized users, confirming that logs are stored locally, and measuring the time to restore functional operation.

This test must occur in the building’s actual environment. A vendor demonstration in a lab does not prove that the local network, controllers, gateways, credentials, and operational procedures will behave correctly at handover.

If a vendor cannot demonstrate the required fallback, the owner has options. The parties can add compensating controls such as redundant gateways or on-site controllers. They can require a vendor-provided emergency access method. They can revise the integration approach. Or they can delay handover until the stated operational requirement has been demonstrated.

That is not an argument against vendors or cloud platforms. It is practical risk management. The building owner is defining what continuity means and requiring evidence that the system meets that requirement.

Lessons From Access and HVAC Outages

One office tower experienced a cloud identity-service failure during a morning tenant shift change. Tenants could not badge in, the lobby became overwhelmed, and contractors could not access critical floors. The building team had a local master-key mode, but it was undocumented and unpracticed. The team eventually activated it through awkward coordination.

After the incident, the building added an on-site authenticated keypad fallback, documented the procedure, and ran a test. The recovery time improved from hours to under 30 minutes. The technology change mattered, but the operational preparation mattered just as much.

In another case, a campus HVAC analytics platform stopped accepting telemetry for 48 hours. Automated cloud-cued setpoint changes did not happen. Temperatures drifted, and occupants complained. The mitigation was not a wholesale replacement of the platform. The team configured local fallback schedules on controllers and added a quarterly blackout exercise to validate the expected behavior. Once the local behavior had been exercised, temperature drift became predictable and escalation responsibilities became clearer.

Both examples demonstrate the same principle: the best fallback is one that is simple, documented, and practiced.

Build a Resilience Program in Five Steps

Owners and operators can begin immediately with a focused checklist.

  1. Inventory every cloud touchpoint. Identify each system that depends on a cloud service, portal, identity platform, telemetry pipeline, or remote control path. Classify it as critical, important, or amenity. Set a 30-day deadline to complete the inventory.
  2. Verify local operation. Confirm what works locally, what does not, and how long the local capability lasts. Test a 24-hour window and determine whether local control can be restored within an acceptable period.
  3. Require cloud-outage acceptance testing. Add cloud failure scenarios to commissioning requirements and require vendors to demonstrate the outcome in the property environment.
  4. Run quarterly exercises. Conduct blackout drills with named roles, communications procedures, and documented outcomes. Exercises turn assumptions into evidence.
  5. Maintain a living playbook. Record escalation contacts, local-log locations, fallback procedures, and tenant-notification templates. Update the playbook after exercises and actual incidents.

What Limited-Staff Teams Should Do First

For a small building team, the minimum viable program is an inventory and one exercise. Identify cloud touchpoints, tag the critical systems, and run a tabletop or simple blackout drill for one critical system.

That one drill can answer the questions that matter most. Does the documented fallback actually work? Does the team know who to call? Are local logs available? Is the vendor contract aligned with the operational need? What needs to change before a real outage occurs?

Cloud technology can be a valuable layer of visibility, support, and convenience. But it should not be the untested foundation of building operations. Design for autonomous operation first. Add cloud convenience on top. Then prove the fallback through acceptance testing, documented playbooks, and regular exercises.

For more practical conversations about infrastructure, building systems, and operational resilience, listen to this episode of Built, Wired & Secured.