GDS Technology — Built, Wired and Secured podcast banner
Watch on YouTube →
Episodes Built
Episode 71

SLA for the Mechanical Room: Setting Practical SLOs for Building Systems

July 7, 2026
Key takeaways
  • Start every outage discussion by asking what breaks if the system goes down.
  • Use SLOs as internal operational targets and keep SLAs focused on contract enforcement.
  • For most building systems, availability, maximum continuous downtime, and response window are enough.
  • Higher uptime targets should be tied to safety, financial impact, or critical operations, not preference.
  • Quarterly tabletop exercises and annual failover testing make building resilience measurable instead of assumed.

Show Notes

When Building Downtime Turns Into a Leadership Decision

This episode of Built, Wired & Secured takes a practical look at how facilities teams, IT leaders, and building owners can define realistic expectations for building-system performance before an outage forces the issue. The conversation opens with a high-pressure scenario: a BAS dashboard goes red, zones are losing control, alarms are stacking up, and a failing UPS is forcing an on-site team to decide what stays powered and what does not.

From there, the discussion focuses on one core idea: the first question in a crisis should be, "What breaks if this goes down?" That question becomes the foundation for triage, resource allocation, and ultimately the Service Level Objectives that guide operations before an emergency ever happens.

SLOs vs. SLAs in Building Operations

A major theme in the episode is the difference between SLOs and SLAs. The distinction is simple but important:

  • An SLA is contract-focused and often tied to penalties.
  • An SLO is an internal operational target that is shared, measurable, and used to guide design and response.
  • SLOs help teams define what level of service is actually required based on the consequences of failure.

That distinction matters because many organizations try to force vendor contracts to carry expectations they have not clearly defined internally. The episode argues that practical SLOs should come first. Once a team knows what acceptable performance looks like, it can build smarter scopes of work, clearer response expectations, and more defensible budgets.

How to Prioritize Systems

The conversation breaks building systems into operational tiers based on impact rather than emotion. Instead of responding to the loudest complaint, teams should rank systems by what happens if they fail.

  • Non-negotiable systems include life safety power, fire detection interfaces, and access control where egress could be affected.
  • High-priority systems include HVAC supporting critical equipment rooms and tenant spaces with immediate operational consequences.
  • Comfort-level systems can tolerate lower targets when the impact is inconvenience rather than safety or major loss.

This framework helps facilities leaders explain why not every system should receive the same uptime target, the same redundancy budget, or the same response commitment.

The Metrics That Actually Matter

Rather than overcomplicating performance management, the episode recommends using just two or three metrics for most systems:

  • Availability percentage
  • Maximum continuous downtime
  • Response window

These metrics create a practical operating model. Availability measures overall reliability. Maximum continuous downtime captures the real pain of a prolonged failure. Response window defines how quickly teams or vendors must engage once a critical issue appears. Together, they provide enough structure to guide action without burying operators in reporting.

Example Targets and Tradeoffs

The episode makes clear that targets should be driven by impact, not by abstract technical ambition. For systems where failure creates safety exposure or major financial loss, the speakers suggest aiming for 99.9% availability or better. For non-critical comfort systems, something closer to 98% may be entirely reasonable.

They also stress that every added "nine" comes with a real cost. Redundancy, infrastructure upgrades, testing discipline, and response capacity all become more expensive as the target rises. That is why leaders should frame the conversation around operational and business consequences rather than a generic push for maximum uptime.

  • Critical control points may justify downtime limits as low as 15 minutes.
  • Amenity or comfort HVAC may tolerate several hours depending on the use case.
  • Critical alarms may require remote diagnosis within 30 minutes.
  • Business-hours on-site dispatch may need a 2-hour target for priority issues.
  • Off-hours expectations can be looser, but they must still be explicit.

Maintenance, Testing, and Honest Failovers

Another strong takeaway is that reliability does not come from design alone. Preventative maintenance, clear labeling, and simple testing are treated as essential. Teams should not assume their failover plans work just because they exist on paper.

The episode warns against two common mistakes: overengineering without discipline, and overconfidence without realistic testing. A single generator or backup path may be acceptable, but only if it is tested under realistic conditions. Tabletop confidence is not the same thing as proven performance under real load.

Hidden Dependencies and Local Fallbacks

Carrier services and hosted platforms are highlighted as a common blind spot. If a hosted BAS or access control system loses connectivity and local controls are not configured to run autonomously, the building can lose effective control even when the underlying equipment is still present.

That means SLOs should account for more than a vendor dashboard promise. They should define both expected carrier performance and what the building does locally when connectivity drops.

How to Measure Without Creating Busywork

The speakers recommend keeping monitoring close to operations and making the reporting sustainable.

  • Use local BAS alarms for direct operational visibility.
  • Add a lightweight cloud dashboard for trend awareness.
  • Maintain simple runbooks for handoffs and escalation.
  • Run quarterly tabletops.
  • Perform at least one annual full transfer test for generators.

Leadership reporting should stay simple: availability percentage, longest continuous downtime, and last test date. If an SLO is missed, teams should produce a root cause note and a corrective plan with a timeline. The point is to improve decisions and guide capital planning, not to create a blame exercise.

Quick Win, Cautionary Tale, and First Steps

The episode closes with two useful examples. In the success case, a medical office tenant had redundant HVAC and a 15-minute failover target for an equipment room, backed by quarterly testing. When a compressor failed, the switchover happened automatically and operations continued without disruption. In the failure case, a campus tried to hold comfort systems to 99.99% availability without funding redundant chillers or realistic response capacity. Teams could not meet the target, and vendor relationships deteriorated.

The practical starting point is clear:

  • Pick two to three metrics for each critical system.
  • Classify systems by safety impact and tenant impact.
  • Schedule quarterly tabletops and at least one full annual failover test.
  • Document who measures, who signs off, and how misses are handled.
  • Use early results to inform capital planning and refine targets over time.

The episode also points listeners to a one-page SLO template and checklist from GDS Technology that can be piloted on a single critical system for one quarter. The bigger message is straightforward: when teams define acceptable downtime, response windows, and fallback behavior in advance, outages become more manageable, decisions become more defensible, and operations become far less reactive.

Deeper dive

Why Building Systems Need Practical SLOs

In commercial buildings, outages rarely begin as abstract technical problems. They begin as operational pressure. A dashboard turns red. Alarms pile up. Someone has to choose what gets attention first. Tenants call. Contractors get pulled in. Leadership wants answers before the root cause is fully understood.

That is the operating environment behind this episode of Built, Wired & Secured, which focuses on a concept more common in IT and software than in facilities: Service Level Objectives, or SLOs. The conversation makes a strong case that building operators should adopt simple, measurable service targets for critical systems so decisions during failures are based on agreed priorities instead of stress, assumptions, or whoever is speaking the loudest.

The central question is refreshingly direct: what breaks if this goes down? Once a team answers that honestly, it can start setting realistic expectations for uptime, downtime, response, redundancy, testing, and ownership.

SLOs Are Not the Same as SLAs

One of the clearest points in the episode is that SLOs and SLAs are not interchangeable. An SLA is generally external and contractual. It often exists between a customer and a provider, and it may include penalties for missed performance. An SLO is different. It is an internal operating target. It defines what a team is trying to achieve, how it will measure that target, and why that target matters.

That difference matters in building operations because many organizations jump straight into vendor expectations before aligning internally on acceptable performance. If the owner, facilities team, and IT group have not agreed on what matters most, vendor scopes and response commitments will often be inconsistent, unrealistic, or too vague to be useful.

Practical SLOs fix that. They turn informal expectations into shared, measurable targets. They also help separate critical risks from comfort-level disruptions, which is essential when budgets are limited and not every system can be engineered to the same standard.

Start With Consequences, Not Complaints

The episode strongly rejects reactive prioritization driven by noise. In a real outage, the loudest caller is not always connected to the highest risk. A tenant discomfort complaint is not automatically equal to a threat to life safety, critical equipment, or controlled egress.

That is why the first question during a failure should be, what breaks if this goes down? The answer shapes both crisis response and SLO design. If a failure creates a safety issue, threatens critical rooms, or causes major operational loss, the system needs a higher target, clearer ownership, and more disciplined testing. If a failure causes inconvenience but not serious risk, the target can be lower and the response model can be simpler.

This consequence-based mindset gives leadership a rational way to prioritize spending, maintenance effort, and contractor accountability. It also creates a more defensible framework for communicating with tenants and stakeholders when not everything can be restored at once.

The Three Metrics That Make SLOs Usable

Another useful theme in the episode is restraint. The speakers do not recommend complicated scorecards with dozens of data points. For most building systems, they argue that two or three metrics are enough.

  • Availability percentage shows overall reliability.
  • Maximum continuous downtime captures how long a system can be out before the impact becomes unacceptable.
  • Response window defines how quickly a team or vendor must engage when a critical issue appears.

That combination is simple enough for facilities teams to use consistently and strong enough to guide real decisions. It also creates a bridge between operations and vendor management, because these are measurable items that can be discussed in service scopes, maintenance plans, and escalation procedures.

Not Every System Deserves the Same Target

A major strength of the episode is its refusal to treat uptime as a one-size-fits-all standard. The speakers divide systems into tiers based on consequence.

At the top are systems where failure is unacceptable: life safety power, fire detection interfaces, and access control where egress is involved. These require the highest availability expectations and the clearest operational ownership.

Next come systems that directly support building operations or critical tenant functions, such as HVAC for equipment rooms and spaces where a loss would cause immediate disruption. Below that are comfort systems, where outages matter but may not justify the same cost or response model.

This tiering is operationally valuable because it prevents teams from making expensive promises in the wrong places. It also helps leadership explain why some systems require redundancy or tighter dispatch windows while others are better served by stronger procedures and realistic restoration plans.

How to Set Numeric Targets Without Guessing

The episode offers straightforward guidance for choosing actual targets. If failure creates safety exposure or major financial consequences, a target around 99.9% availability or higher may be justified. For non-critical comfort systems, something lower, such as 98%, may be fully acceptable depending on building use and stakeholder expectations.

The same principle applies to maximum continuous downtime. A critical control point may warrant a limit measured in minutes. An amenity HVAC issue may be acceptable for hours. The lesson is not to copy generic standards. It is to tie the target to the consequence.

That framing also helps when cost becomes a concern. Rather than arguing over technical decimals, leaders can ask a better question: what downstream risk does this investment reduce? If higher availability meaningfully lowers tenant churn, protects compliance, or prevents equipment loss, the spend is easier to justify. If not, the better answer may be a lower target paired with stronger procedures and response planning.

Reliability Depends on Maintenance and Testing

One of the most practical parts of the conversation is the emphasis on disciplined maintenance. The speakers make it clear that reliability is not only a design issue. Preventative maintenance, labeling, runbooks, and routine testing all matter because they determine whether the system behaves the way the team thinks it will behave during a real event.

The conversation also pushes back on false confidence. A backup path that has never been tested under realistic load is not proven resilience. A tabletop exercise is useful, but it does not replace a real transfer test. If a building depends on a single generator, that dependency has to be tested honestly. If a backup HVAC path is supposed to protect an equipment room, the team needs evidence that the switchover works inside the expected time window.

This is an important business point as much as a technical one. Unverified assumptions create avoidable surprises, and surprises are expensive.

Watch the Hidden Dependencies

The discussion around carrier outages highlights a problem many buildings underestimate. A hosted BAS or access control platform may appear available until connectivity drops. If local controls are not configured to operate autonomously, the building can effectively lose control even though the equipment remains installed and powered.

That means good SLOs should include both service availability and fallback behavior. It is not enough to ask whether the carrier stayed up. Teams also need to know what the building does locally when the network path fails, who owns that behavior, and whether it has actually been tested.

Use SLOs to Drive Better Operations, Not Blame

Measurement only helps if it stays close to operations. The episode recommends a sustainable rhythm: local BAS alarms, a lightweight cloud dashboard for trends, simple runbooks for handoffs, quarterly tabletop exercises, and an annual full transfer test for generators. Leadership reporting can then stay concise: availability percentage, longest continuous downtime, and last test date.

That reporting model matters because it keeps SLOs from becoming another ignored spreadsheet. It also creates a useful rule when targets are missed: document the root cause, assign a corrective plan, and attach a timeline. In other words, use SLO misses to improve the environment and inform capital planning, not to punish teams for working inside known constraints.

A Practical Starting Point

The episode closes with a grounded first-step framework. Pick two or three metrics for each critical system. Classify systems by safety and tenant impact. Define who measures performance, who signs off, and how misses are handled. Schedule quarterly tabletops and at least one full annual failover test. Then start with one critical system for one quarter, gather real data, and refine the targets based on evidence.

That is a smart way to start because it avoids both paralysis and overreach. Teams do not need a perfect enterprise-wide program on day one. They need a repeatable method that turns vague expectations into measurable operating standards.

For owners, operators, and facilities leaders, that is the real value of the episode. It is not just about uptime math. It is about reducing surprises, clarifying ownership, and making building-system decisions that hold up under pressure. If you want a better way to align facilities, IT, and vendors around real-world performance, this conversation is worth a listen.