Show Notes
When Building Data Stops Helping
Modern buildings generate a constant stream of telemetry: access control alerts, HVAC faults, UPS status flags, carrier and network health pings, and vendor notifications. In this episode of Built, Wired & Secured, Alex Morgan and Michael Harrington explain what happens when that data stops serving operations and starts creating friction. They call that problem telemetry debt: the buildup of noisy, unmanaged, poorly owned signals that slow teams down instead of helping them respond faster.
The episode opens with a scenario many property and facilities teams will recognize: multiple alarms hit at once, tenants start calling, and nobody is sure which issue matters most. The central point is simple. More telemetry does not automatically produce better decisions. Without signal quality, clear ownership, and a practical triage process, data becomes operational drag.
What Causes Telemetry Debt
Michael points to several causes that stack on top of each other:
- Vendor default alerts are often too verbose and treat every sensor change like a meaningful warning.
- Retention settings keep months of raw logs that nobody actively reviews.
- Ownership is unclear across IT, facilities, and outside vendors.
- There is often no agreed triage workflow for what happens when an alert arrives.
One example from the episode is a badge reader reporting low RSSI every five minutes. On its own, that may be a transient issue. Across hundreds of readers, it becomes a flood of warnings that hides genuinely important failures.
Start With Risk, Not Convenience
A key theme in the conversation is that teams should not think about telemetry cleanup as simply turning alarms off. The better framing is making alarms useful. Alex and Michael emphasize that signal decisions should start with risk. If missing a signal would create tenant impact, a safety problem, or a prolonged outage, that signal should remain intact. Lower-value signals can be rolled up, sampled, compressed, or delayed.
That approach creates a more defensible operational model. Rather than arguing about whether there is too much data in the abstract, teams can ask a clearer question: what breaks if this goes down? That question helps prioritize which alerts deserve immediate attention and which ones should be handled with lighter retention or less aggressive alerting.
The Low-Effort Fixes That Move Fast
The episode focuses on practical changes teams can make without launching a long transformation project. Three actions stand out:
- Set sensible alert thresholds so teams can tell the difference between a blip and a trend.
- Create an ownership matrix that names a primary owner, a secondary owner, and the vendor for each alert category.
- Use tiered retention so critical systems keep full-fidelity data while lower-risk systems are compressed or sampled.
These are not heavy technical changes. They are operating decisions that improve signal quality and reduce wasted time.
A One-Page Triage Model
One of the most actionable parts of the episode is the discussion of first-response triage. Michael breaks it into three steps:
- Verify: confirm whether the alert can be reproduced or validated with a single supporting data point.
- Identify ownership: know who the primary contact is and who is on call next.
- Contain: take the smallest action that reduces tenant impact while investigation continues.
If containment does not work, the next move is to escalate to the vendor with a packaged set of logs. That reduces unnecessary back-and-forth and keeps handoffs from becoming a black hole.
Why Tiered Retention Matters
The conversation also addresses a common objection: teams often resist pruning telemetry because they worry about losing data for post-incident forensics. Michael argues that this concern is real, but it does not justify identical retention across every system. Tiered retention offers a more balanced answer. Systems with high forensic value should keep detailed logs. Lower-risk devices can be sampled or compressed without compromising meaningful investigation later.
The practical lesson is that cleanup does not mean ignorance. It means aligning data depth with business value and operational risk.
What Success Looks Like
Rather than measuring success by raw acknowledgment speed, Michael recommends tracking mean time to meaningful action. Teams should also watch repeated alerts per device and tenant-impact events, such as how often tenant calls coincide with alert storms. Those measurements reveal whether cleanup is actually improving decision-making.
Two examples from the episode make the value concrete. In one anonymous campus pilot, roughly 1,200 daily alerts across access control and HVAC were addressed through an ownership matrix, adjusted thresholds, and reduced debug retention for non-critical devices. Alerts fell by about 70 percent, and mean time to meaningful action dropped from four hours to under 90 minutes. In another example, a managed office portfolio reduced overnight noise by adding a delay-and-confirm rule for transient sensors. If an alarm cleared within five minutes, no ticket was created. That cut midnight wake-up calls in half without missing outages.
Three Actions to Start This Week
For teams ready to act, the episode closes with a simple starting checklist:
- Run a seven-day audit of alert volume and map who touches each alert.
- Define retention windows for critical and non-critical telemetry, then document the rationale.
- Create a one-page triage playbook and test it on one building or one system first.
The bigger takeaway is that telemetry cleanup is not about reducing visibility. It is about restoring clarity, speeding containment, and protecting tenant experience. Buildings already produce the data. The real operational advantage comes from deciding which signals matter, who owns them, and what happens next.
Telemetry Debt Is a Building Operations Problem, Not Just a Data Problem
Buildings are producing more telemetry than ever before. Access control systems generate forced-door alerts and reader health messages. HVAC platforms surface fault alarms and performance anomalies. UPS systems issue aging flags and power warnings. Vendors add their own health pings and log streams. On paper, that sounds like progress. More visibility should mean better operations.
But as discussed in this episode of Built, Wired & Secured, that is not what many teams experience in practice. Instead, they face what Alex Morgan and Michael Harrington call telemetry debt: a buildup of noisy, poorly managed signals that creates confusion, slows response, and buries the events that actually matter.
The issue is not a lack of data. In many environments, there is plenty of it. The issue is that the data is not organized around meaningful action. When multiple alerts fire at once, tenants call before teams can get aligned, and internal responders cannot agree on what to handle first, the building has crossed from observability into operational drag.
Why More Telemetry Can Create Less Clarity
One of the most useful points in the episode is that telemetry debt usually comes from both technology defaults and organizational choices. Vendors often ship systems with highly verbose alerting because that showcases capability. Every sensor fluctuation, weak signal, or transient warning gets exposed. That may look comprehensive during procurement, but in a live building it can overwhelm the people responsible for acting on the information.
Michael gives a strong example: a badge reader reporting low RSSI every five minutes. In isolation, that warning might not matter. Across a few hundred devices, it becomes a constant stream of noise. Most of those events may be transient and harmless, yet they crowd out the few alerts that represent real service risk.
The problem gets worse when teams keep months of raw logs they never review and fail to define clear ownership. Is IT supposed to handle network pings? Does facilities own HVAC chimes? Is the vendor responsible for everything because the platform came from them? Without clear answers, alerts bounce between teams and lose urgency. By the time someone acts, the tenant experience may already be affected.
Start With the Business Question: What Breaks If This Goes Down?
Telemetry cleanup often gets framed the wrong way. People hear it as a proposal to reduce visibility or silence alarms for convenience. The episode argues for a much better framing: make alarms useful.
The first question Michael asks is, “What breaks if this goes down?” That question immediately ties telemetry decisions to business impact. If a missed signal could create a safety issue, a tenant-impacting outage, or a long recovery, that signal should remain prominent. If the signal is low-risk, repetitive, or rarely actionable, it can be sampled, delayed, rolled up, or compressed.
This is an important distinction. The goal is not to eliminate telemetry. The goal is to make sure teams can separate meaningful risk from background chatter. When every signal looks urgent, nothing is truly urgent.
The Practical Tradeoffs That Improve Operations Fast
The conversation stays refreshingly practical. Rather than prescribing a massive transformation, Alex and Michael focus on low-effort changes that can reduce noise quickly.
The first is better thresholds. Teams need to distinguish between a temporary blip and a real trend. If thresholds are too sensitive, every short-lived issue becomes a page, a ticket, or a wake-up call. If thresholds are too loose, real failures go unnoticed for too long. The right thresholding model gives teams enough warning to act without treating every transient event as an incident.
The second is an ownership matrix. This may sound administrative, but it has a direct operational payoff. Every alert category should map to a primary owner, a secondary owner, and the vendor. Just as important, the primary owner should have the first three actions defined in advance. When an alert arrives, people should not have to debate responsibility in real time.
The third is tiered retention. Not every stream of telemetry deserves the same storage depth or fidelity. Critical systems with high forensic value may justify detailed retention. Lower-risk devices often do not. Sampling or compressing lower-value telemetry reduces clutter without giving up meaningful operational insight.
A Better First Response: Verify, Identify Ownership, Contain
One of the strongest operational takeaways from the episode is the three-step triage model Michael outlines.
First, verify. Can the alert be reproduced, or can it be confirmed through one supporting data point? This step prevents teams from burning time on false positives or one-off noise.
Second, identify ownership. Who is the primary responder? Who is next on call? If that answer is unclear, the problem is not just the alert. It is the operating model behind it.
Third, contain. Take the smallest action that reduces tenant impact while the investigation continues. Containment is about preserving service and buying time, not solving every root cause immediately.
If containment does not resolve the issue, escalation to the vendor should include a packaged set of logs. That is a subtle but important operational improvement. Instead of tossing the issue over the wall and waiting, teams hand off enough context to move the case forward.
Retention Does Not Need to Be Uniform to Support Forensics
Another useful point from the discussion is the challenge around post-incident forensics. Teams often defend broad retention on the grounds that they may need the data later. That concern is valid, but it can become a blanket excuse for retaining everything at the same level forever.
The episode’s answer is tiered retention. Systems with meaningful forensic value should retain detailed data. Low-risk devices do not need identical treatment. This approach respects the realities of incident review while still reducing telemetry bloat. It also creates a clearer governance model because the retention decision is intentional and documented rather than accidental.
What Should Teams Measure?
If leadership is going to approve a telemetry cleanup effort, they need proof that it worked. Michael recommends focusing on mean time to meaningful action rather than just mean time to acknowledge. That shift matters. A fast acknowledgment means little if the team still spends hours figuring out what to do next.
He also suggests tracking repeated alerts per device and tenant-impact events, particularly how often tenant calls correlate with alert storms. Those indicators connect telemetry quality to the real-world operating experience of both staff and occupants.
When cleanup is working, the outcome is not just a quieter inbox. Teams should see fewer repetitive alerts, faster containment, and fewer tenant escalations.
Two Examples That Make the Case
The episode includes two concrete examples that show why this matters.
In one anonymous campus environment, operations teams were dealing with around 1,200 daily alerts across access control and HVAC. Over a two-week pilot, they introduced an ownership matrix, adjusted thresholds, and reduced debug retention for non-critical devices. The result was roughly a 70 percent drop in alerts and a reduction in mean time to meaningful action from four hours to under 90 minutes. Tenants called less, and the operations team had more capacity to focus on the events that actually needed intervention.
In another example, a managed office portfolio was dealing with vendors waking facilities teams at 2:00 a.m. for transient sensor events. The solution was a delay-and-confirm rule: if an alarm cleared within five minutes, no ticket was created. That cut midnight wake-up calls in half without increasing missed outages. The benefit was operational and human. Staff well-being improved, and tenants were no longer escalating to the property manager in the middle of the night.
Three Actions Teams Can Start This Week
The episode closes with a practical three-step starting point that building leaders can put into motion quickly.
- Run a seven-day audit of alert volume and identify who touches each alert.
- Agree on retention windows for critical versus non-critical telemetry and document why those tiers exist.
- Create a one-page triage playbook with the first three response steps, then test it on a single building or system.
That is an important operational lesson in itself. Cleanup efforts do not have to begin at portfolio scale. A focused pilot creates evidence, lowers resistance, and gives leadership something measurable to review.
The Real Goal: Better Tenant Experience and Faster Response
The strongest closing point in the episode is that teams should avoid talking about telemetry cleanup as alarm reduction for convenience. The better message is usefulness. The point is to improve uptime, reduce preventable confusion, and protect tenant experience.
Buildings already have the data. The challenge is deciding which signals matter, who owns them, and what the first response should be. That is where telemetry debt becomes solvable.
If this episode reflects problems your team sees every week, it is worth listening through the full conversation and using the checklist and pilot template referenced in the episode as a starting point. The fastest operational wins often come not from adding more tools, but from making existing signals clearer, more accountable, and more actionable.