GDS Technology — Built, Wired and Secured podcast banner
Watch on YouTube →
SLA for the Mechanical Room: Setting Practical SLOs for Building Systems
Episodes General
Episode 71

SLA for the Mechanical Room: Setting Practical SLOs for Building Systems

July 7, 2026
Key takeaways
  • Start every building-system SLO by asking what breaks if the system goes down and prioritize by consequence, not noise.
  • For most systems, three metrics are enough: availability percentage, maximum continuous downtime, and response window.
  • Life safety, critical power, fire interfaces, and egress-related access control need the highest targets and clearest ownership.
  • Carrier-dependent systems should include local fallback behavior in the SLO, not just carrier uptime.
  • Use SLO misses to drive root-cause review, corrective action, and capital planning rather than blame.

Show Notes

Why building-system expectations break down in the real world

This episode of Built, Wired & Secured opens with a high-pressure scenario that feels familiar to anyone responsible for a commercial facility: the BAS dashboard turns red, zones start losing control, alarms stack up, a UPS begins tripping breakers, tenants are calling, and someone on site has to decide what stays powered and what does not. The conversation quickly makes the central point of the episode clear. In those moments, operations cannot rely on vague expectations or whoever is shouting the loudest. They need simple, measurable service level objectives that define what matters most before the outage happens.

The discussion centers on a practical framework for setting SLOs for building systems including HVAC, BAS, power, carrier-dependent services, and physical security. The focus is not theory for its own sake. It is how facilities, IT, operators, and building owners can make more defensible decisions about maintenance, redundancy, vendor response, and capital planning.

SLOs vs. SLAs: the distinction that changes how teams operate

One of the most useful parts of the episode is the clear distinction between an SLA and an SLO. The speakers explain that an SLA is contract-focused and often tied to penalties, while an SLO is an internal operational target. That distinction matters because teams often try to force contract language onto systems they do not fully control.

  • An SLA is about contractual accountability.
  • An SLO is about an internal, shared, measurable operational target.
  • SLOs should guide design, response planning, maintenance, testing, and escalation.
  • Teams should avoid turning internal SLOs into punitive SLAs until they actually control the levers required to meet them.

That framing keeps the conversation grounded. Instead of debating abstract uptime numbers, teams can define what level of performance is necessary based on operational impact.

Start with consequence, not noise

A recurring theme throughout the episode is triage by consequence. The first question raised is simple: what breaks if this goes down? According to the discussion, that one question should drive sequence and resource allocation during an incident and should also become the anchor for an SLO.

The speakers draw a hard line between discomfort and unacceptable risk. If failure threatens life safety, critical equipment rooms, or egress-related access control, those systems get protected first. Tenant discomfort matters, but it is not the same as risking safety or damaging critical infrastructure. That kind of prioritization makes SLOs more than reporting metrics. It turns them into decision tools.

The metrics that matter most

The episode strongly argues for simplicity. Rather than drowning operations teams in dashboards and excessive KPIs, the recommendation is to keep most systems to two or three core measures:

  • Availability percentage
  • Maximum continuous downtime
  • Response window

Those three metrics are presented as enough to cover safety, tenant impact, and operational capability for most building systems. They are also practical enough to track consistently.

The speakers then map those metrics to system types. Non-negotiable systems such as life safety power, fire detection interfaces, and core access control need the highest availability targets and the clearest ownership. High-priority HVAC tied to critical equipment rooms or spaces with immediate operational consequences follows next. Comfort-level systems can usually tolerate lower targets.

How to choose the right target numbers

The episode offers concrete examples for setting SLO values without pretending every building needs the same thresholds. If a failure risks safety or major financial loss, the guidance is to aim for 99.9% availability or better. For non-critical comfort systems, 98% may be acceptable. Maximum continuous downtime should also vary by impact. A critical control point may only tolerate 15 minutes, while amenity HVAC might tolerate several hours.

Just as important, the conversation does not oversell high-availability design. The tradeoff is stated directly: cost rises fast as teams chase each additional nine. That makes the downstream impact question essential.

  • If higher availability prevents tenant churn, compliance problems, or major operational disruption, it is a risk mitigation investment.
  • If the issue is mainly comfort, teams may be better off setting a lower target and investing in response playbooks instead of expensive redundancy.

Response windows, carriers, and hidden dependencies

The episode also highlights that availability alone is not enough. Teams need explicit response windows tied to incident severity and hours of operation. One practical model shared is:

  • 30 minutes for remote diagnosis on critical alarms
  • 2 hours for on-site dispatch during business hours
  • Looser off-hours targets, but only if they are clearly documented

This becomes especially important when carrier services or hosted platforms are involved. One speaker points out that outages often expose hidden dependencies. If a hosted BAS or access control vendor loses connectivity and local controls are not configured to run autonomously, the building effectively loses control. That means a building-system SLO should cover both carrier availability and fallback behavior when connectivity is lost.

Basics before overengineering

A valuable tension in the episode comes from the debate over redundancy versus operational discipline. On one side, the speakers argue that many mid-market buildings do not need top-tier redundancy everywhere. Disciplined maintenance, clear labeling, local backups, and simple tests often deliver more reliability per dollar than overengineered solutions.

At the same time, they caution that exercises cannot compensate for inadequate design. If a site depends on a single generator, it needs to be tested under realistic load. Tabletop confidence is not enough if the actual system cannot carry real demand. The takeaway is balanced and practical: reliability depends on both sound design and honest testing.

Monitoring, ownership, and an operating rhythm teams can sustain

The episode avoids suggesting an unrealistic monitoring stack. Instead, it recommends keeping measurement close to operations. Local BAS alarms, a lightweight cloud dashboard for trend visibility, and simple runbooks are presented as a sustainable model.

The operating rhythm recommended in the episode includes:

  • Quarterly tabletops
  • An annual full transfer test for generators
  • Monthly scorecard reporting to facilities leadership and IT

That scorecard should remain simple:

  • Availability percentage
  • Longest continuous downtime
  • Last test date

When an SLO is missed, the action should be straightforward: document root cause, define a corrective plan, and assign a timeline. The point is not blame. The point is using SLO data to inform future capital decisions and operational improvements.

Real examples: what works and what backfires

Two examples bring the framework to life. In the positive case, a medical office tenant had redundant HVAC and a 15-minute failover objective for an equipment room, supported by quarterly testing. When a compressor failed, automatic switchover worked and the failure passed unnoticed. That is the kind of reliability operators want: success that feels uneventful because the system performed the way it was designed and tested to perform.

The cautionary example goes the other direction. A campus was pushed toward 99.99% availability for comfort systems without funding redundant chillers, while response windows were set tighter than teams could realistically meet. The result was predictable: expectations became punitive, vendors were frustrated, and the targets were not credible. The lesson is clear. Teams should not promise performance they do not have the infrastructure, staffing, or vendor structure to support.

The fastest way to get started

The episode closes with a practical three-step starting point for any organization that wants to launch an SLO program without overcomplicating it:

  • Pick two to three metrics per critical system.
  • Classify systems by safety and tenant impact so priorities are documented and shared.
  • Schedule simple tests including quarterly tabletops and at least one full failover annually.

The speakers also stress the need to document ownership: who measures, who signs off, and how misses are handled. From there, teams can collect data, refine targets, and use the results to guide capital planning with evidence instead of guesswork.

The episode ends with an invitation to pilot a one-page SLO template and checklist in a single building for one quarter. That advice captures the overall tone of the conversation: keep it practical, keep it measurable, and build credibility through repeatable operating discipline.

Deeper dive

Building-system reliability needs clearer expectations, not louder arguments

When building operations go sideways, the first failure is often not the equipment. It is the lack of a shared definition of what “acceptable” was supposed to mean in the first place.

That is the central lesson from this episode of Built, Wired & Secured, which explores how facilities teams, IT, building owners, and vendors can use simple service level objectives to turn vague expectations into measurable operating targets. The conversation stays rooted in real-world building systems: HVAC, BAS, power, carrier-dependent services, and physical security.

The episode opens with a scenario that instantly makes the stakes feel real. It is the middle of the day. The BAS dashboard goes red. Zones are losing control. Alarms are piling up. A failing UPS is tripping breakers. Tenants are calling. Someone on site has to decide which systems stay powered and which ones do not.

In that moment, teams do not have time to debate philosophy. They need a framework that tells them what matters first, how long a system can be down, what response is required, and who owns the next step. That is where practical SLOs come in.

Why SLOs matter more than most facilities teams realize

One of the most useful ideas in the episode is the distinction between an SLA and an SLO. It sounds minor, but operationally it changes everything.

An SLA is a contract tool. It is usually tied to obligations, penalties, and vendor accountability. An SLO is an internal operational target. It is shared, measurable, and designed to guide how a building is engineered, maintained, monitored, and supported.

That distinction matters because many organizations create frustration by demanding SLA-style outcomes from systems that have never been designed, funded, or staffed to meet those expectations. When teams skip the internal discipline of setting realistic SLOs first, outages become emotional, escalations become political, and vendors end up being blamed for structural gaps the building owner never addressed.

Practical SLOs give everyone the same operating language. They make tradeoffs visible before an incident. They also create a clearer connection between system design and business impact.

Start with consequence, not complaints

The episode keeps returning to one simple question: what breaks if this goes down?

That question is useful because it cuts through noise. During an outage, the loudest stakeholder is not always the one facing the biggest risk. Tenant discomfort may matter, but it is not equivalent to a life safety issue, a critical equipment-room failure, or an egress problem tied to access control.

That means the right first move is triage by consequence. Systems tied to life safety, fire detection interfaces, critical power, or egress-related access control belong at the top of the hierarchy. HVAC tied to critical equipment rooms comes next. General comfort systems can usually tolerate more downtime and looser response expectations.

This is exactly where SLOs add value. They force operators and owners to define, in advance, what level of service must be guaranteed because the consequence of failure is unacceptable. That makes real-time decision-making faster and more defensible.

The simplest metrics are often the most useful

The discussion pushes back against the idea that facilities teams need complex reliability frameworks to improve operations. For most building systems, the speakers recommend tracking only two or three core metrics:

  • Availability percentage
  • Maximum continuous downtime
  • Response window

That is a strong model because it balances discipline with usability. Availability percentage tells you how often the system was up. Maximum continuous downtime tells you how painful a single failure can be. Response window ties the system target to human action, whether that means remote diagnosis, dispatch, or escalation.

Most importantly, those metrics are simple enough to survive contact with reality. If a scorecard requires too much interpretation or too many exceptions, it will not be used consistently. A facilities program that lasts usually depends on measures operators can track without a data science project.

How to set targets without pretending every building is the same

The episode gives practical numeric examples while still respecting building-to-building differences. If failure risks safety or major financial loss, the recommended target is 99.9% availability or better. For non-critical comfort systems, 98% may be enough. Critical control points may deserve a maximum continuous downtime of 15 minutes. Amenity HVAC can often tolerate several hours.

Those examples matter because they show how targets should be built from impact, not copied from a generic standard.

The conversation also handles a reality many owners avoid: each additional nine gets expensive fast. Chasing higher availability numbers without understanding the consequence of failure is a recipe for overspending in the wrong places.

A better decision model is straightforward. If a higher availability target protects against tenant churn, compliance exposure, or meaningful operational loss, it may be justified as risk mitigation. If the system mainly affects comfort, the smarter move may be a lower target paired with better response playbooks and clearer communications.

That is not a compromise. It is prioritization.

Carrier dependencies and local fallback deserve more attention

One of the more important operational points in the episode is that buildings increasingly depend on carrier-connected services and hosted platforms. BAS tools, access control systems, and other building operations can appear healthy until connectivity fails. If local controls are not configured to operate autonomously, the building can lose practical control even when the physical systems themselves are still in place.

That means a serious SLO cannot stop at carrier uptime alone. It needs to define fallback behavior. What happens locally if the hosted system becomes unreachable? What remains controllable on site? How long can that degraded mode be tolerated? Who verifies that fallback actually works?

This is a good example of why building-system reliability is no longer just a facilities issue or just an IT issue. It lives in the overlap.

Why disciplined basics often outperform expensive complexity

The episode strikes a useful balance on redundancy. It does not argue that every mid-market building needs top-tier infrastructure everywhere. In fact, one of the clearest points made is that disciplined maintenance, clear labeling, local backups, and straightforward testing often buy more reliability per dollar than overengineered solutions.

That is an important message for owners trying to allocate limited capital. Too many reliability conversations jump straight to hardware duplication while ignoring the basics that make systems survivable in the field.

At the same time, the speakers warn against confusing preparedness with wishful thinking. Tabletop exercises are valuable, but they do not replace real design capability. If a site relies on a single generator, the team needs to verify failover under realistic load conditions. A test that never reflects actual operating conditions can create false confidence.

The takeaway is practical: design and practice are both required. You cannot talk your way around weak infrastructure, and you cannot buy your way around poor operating discipline.

Ownership, monitoring, and a rhythm teams will actually maintain

Another strength of the episode is that it recommends an operating model most organizations can sustain. Instead of a heavy monitoring stack, the speakers advocate for keeping measurement close to operations through local BAS alarms, lightweight cloud dashboards for trend visibility, and simple runbooks that clarify handoffs.

The recurring rhythm they recommend is equally practical:

  • Quarterly tabletop exercises
  • An annual full transfer test for generators
  • A monthly scorecard reviewed by facilities leadership and IT

That scorecard should be short enough to act on:

  • Availability percentage
  • Longest continuous downtime
  • Last test date

If an SLO is missed, the response should not be finger-pointing. It should be root cause, corrective plan, and timeline. This is where mature building operations separate themselves from reactive ones. Misses should improve capex planning and operating decisions, not just generate blame.

What success and failure look like in practice

The episode gives two examples worth carrying forward. In the success case, a medical office tenant had redundant HVAC and a 15-minute failover target for an equipment room backed by quarterly tests. When a compressor failed, automatic switchover worked and the failure was essentially invisible to occupants. That is what well-designed reliability looks like: not drama, but silence.

The failure example is just as instructive. A campus pursued 99.99% availability for comfort systems without funding redundant chillers, while also setting response windows teams could not realistically meet. The result was predictable. The targets were not credible, vendors became frustrated, and the entire framework created tension instead of resilience.

The lesson is simple. Do not set targets that your design, staffing, and vendor structure cannot support.

A practical starting point for owners and operators

If there is one reason this episode stands out, it is that it gives a usable starting point instead of just describing the problem. The recommended first steps are clear:

  • Choose two to three metrics for each critical system.
  • Classify systems by safety and tenant impact.
  • Run simple, repeatable tests including quarterly tabletops and at least one annual full failover.

From there, document ownership. Decide who measures performance, who signs off on results, and how misses are handled. Then collect data for a quarter and refine the targets based on evidence.

For commercial buildings, that approach has real strategic value. It improves outage response, sharpens vendor expectations, supports maintenance prioritization, and creates a more defensible path for capital planning. In other words, it helps owners and operators move from reactive operations to a technology and facilities strategy they can actually govern.

If this episode resonates, it is worth listening in full. The conversation does a strong job of translating reliability concepts into building-language that operators, IT teams, and ownership groups can all use.