Show Notes
Why access control failures become operational problems fast
This episode opens with a familiar building operations nightmare: it is 8:42 on a Tuesday, the lobby doors are behaving badly, and the entire property feels the impact immediately. The conversation frames access control issues not as isolated hardware glitches, but as business continuity, tenant experience, safety, and liability problems. Alex Morgan is joined by Michael Harrington and James Rogers to break down what actually fails first, what human behaviors make incidents worse, and what teams can do this week to make access systems more resilient without creating new security gaps.
One of the clearest takeaways is that power and connectivity still drive many access control failures. When a building loses power, when UPS capacity is too small, or when controllers depend too heavily on cloud communication, doors stop behaving the way operators expect. The panel stresses that organizations need a door-by-door inventory that answers practical questions ahead of time: which doors lock, which doors unlock, and which ones rely on live backend communication. Without that inventory, teams are guessing during the worst possible moment.
The failure modes that hide until the wrong day
Beyond power and network issues, the episode highlights quieter failure modes that often stay hidden until a patch, a weekend outage, or a busy morning exposes them. These include battery exhaustion on wireless locks, expired certificates, and corrupted local credential caches. What makes these issues so dangerous is that the system can appear healthy right up until the moment it is under stress.
- Battery-backed devices may not fail gradually in a visible way if remote reporting is missing or ignored.
- Certificate and cache issues can stay unnoticed until connectivity changes or routine maintenance occurs.
- Controllers that rely on live services can create bottlenecks when carriers or backend platforms go down.
- Poor alert routing can turn a technical issue into a delayed operational response.
The speakers also spend time on what people do when the system fails. Queues form, doors get propped open, and credentials get shared. Those short-term workarounds may restore movement, but they create larger security and audit problems. That framing matters because it shifts the conversation from devices alone to the full operating environment around them.
Mechanical overrides versus digital resilience
A central tension in the episode is the tradeoff between mechanical overrides and digital backup methods. Michael Harrington is cautious about mechanical keys because they break auditability. Once a key is handed out, the organization often loses the ability to answer who used it and when. He points out that human behavior makes the problem worse, with keys ending up on lanyards or in a drawer marked master.
James Rogers agrees that the risk is real, but argues that a controlled mechanical option can still be the safest immediate response during a power outage. The key is discipline. He recommends dual custody, signed logs, or sealed key boxes with manager access so overrides remain temporary, controlled, and documented. Michael agrees with that approach when the mechanical path is treated as an emergency-only tool rather than a daily convenience.
That distinction is one of the most useful themes in the conversation: resilience is not just about adding more ways in. It is about adding the right fallback methods with clear controls so the backup does not become a permanent vulnerability.
Why local caching needs risk-based rules
The discussion then shifts to local caching versus real-time credential checks. James gives the short version: local caching lets a reader store permissions and keep operating when the server is unavailable. That improves resilience during short outages and keeps people moving. The tradeoff is stale data. If a badge should be revoked or a temporary suspension needs to take effect, a cached reader may not know until it reconnects.
Michael pushes the point further by emphasizing that a lab and a break room should not follow the same caching policy. Sensitive zones need shorter cache windows and stronger physical fallback controls. General office spaces can tolerate longer windows. The result is a tiered model with different cache durations by zone and clearly documented risk tolerances.
- General office areas may support longer cache windows to preserve movement during short outages.
- Higher-security areas need shorter cache windows to reduce stale permission risk.
- Fallback methods should reflect the sensitivity of each space, not a one-size-fits-all policy.
- Documented zone-based tolerances help facilities, security, and IT make consistent decisions.
Power strategy, battery strategy, and alert ownership
When the conversation turns practical, the speakers recommend tiered power planning. For perimeter doors and critical controllers, James recommends UPS coverage in the four-to-eight-hour range or enough runtime to bridge to generator transfer. For interior doors, battery-backed locks with remote battery reporting and scheduled replacement cycles can reduce surprises.
Just as important, the panel warns listeners not to trust vendor defaults. A battery alert that never creates a ticket, or a device warning that lands in the wrong queue, is not a resilience strategy. Someone has to own follow-up. That ownership point repeats throughout the episode and connects the technical system to real operating accountability.
A 20-minute tabletop teams can run this week
One of the most actionable parts of the episode is James Rogers' short tabletop exercise for a controller that loses cloud connectivity during morning ingress. The point is not to build a giant emergency program. It is to practice one realistic scenario with the people who will actually be on the hook.
- Minute 0: facilities checks UPS and power conditions.
- Minutes 1 to 5: security tests whether cached credentials still allow access.
- Minutes 6 to 10: IT checks carrier and backend service status.
- Minutes 11 to 15: the team practices the approved temporary access method, such as a controlled mechanical override with sign-in.
- Minutes 16 to 20: debrief, document what happened, and log changes to the SOP.
The speakers add one more operational rule that matters: if the issue is not resolved within 30 minutes, escalate to building management and send a tenant notification. That communication step is framed as a way to reduce surprise, reduce frustration, and prevent people from propping doors open while they wait for answers.
The lean spare kit and the role of frontline staff
For day-one preparedness, the episode recommends a simple indexed on-site spare kit. Michael calls out spare readers, a controller module, standard batteries, and documented firmware images as a lean set that can make a real difference. He also stresses including vendor responsibilities and SLAs in the documentation so teams are not arguing over ownership during an incident.
James adds that frontline staff must be empowered and trained. If a receptionist has to call multiple people just to use an approved override, they will likely improvise. Practiced, documented, easy-to-execute procedures are more resilient than policies that look good on paper but fail under pressure.
Three actions to take right now
By the end of the episode, the recommendations are concrete. Michael says listeners should document fail modes for critical doors, confirm that device alerts land in a monitored ticketing queue with a real owner, and prepare a 30-minute escalation script plus a tenant notification template. James adds that teams should run a 20-minute tabletop this week, build a small spare kit, and begin battery reporting so replacement cadence is guided by real telemetry instead of assumptions.
The closing message ties physical resilience and cybersecurity together. A backup method that introduces a network exposure is not a win. The goal is to keep people moving, keep buildings secure, and make failures boring through planning, testing, and disciplined operations.
When access control fails, the problem is bigger than the door
Most organizations do not think deeply about access control resilience until a bad morning forces the issue. In this episode of Built, Wired & Secured, Alex Morgan sits down with Michael Harrington and James Rogers to talk through what happens when electronic access systems stop behaving the way operators expect. The topic is physical access, but the conversation quickly shows why it also belongs in the worlds of operations, tenant experience, safety, and cybersecurity.
The scenario that opens the episode is intentionally simple and uncomfortable: it is 8:42 on a Tuesday, and the lobby doors either will not stop letting people in or will not let anyone out. That kind of failure creates pressure immediately. People line up. Guards get overwhelmed. Tenants get frustrated. Someone props a door. Someone shares credentials. A technical issue becomes a human issue in minutes.
That is what makes this episode valuable. It does not treat resilience as an abstract design goal. It treats it as a practical operating discipline that reduces drama when real systems misbehave.
What usually fails first
According to the discussion, power and connectivity lead the list. If a building loses incoming power, if the UPS is undersized, or if controllers depend too heavily on live cloud communication, access systems stop acting predictably. The panel makes a basic but important recommendation: inventory your critical doors before a failure happens.
That inventory should answer operational questions, not just asset-management questions. Which doors fail locked? Which fail unlocked? Which depend on a live server? Which ones can continue functioning with local permissions? Without those answers documented ahead of time, a response team is making decisions under pressure instead of following a plan.
The episode also calls out quieter failure modes that are easy to ignore:
- Battery exhaustion on wireless locks
- Expired certificates
- Corrupted local caches
- Device alerts that never reach a monitored queue
These are the kinds of issues that stay hidden during normal operation and appear only when a patch, outage, or high-traffic event exposes the gap. The lesson is straightforward: resilience failures are often maintenance failures that have not been discovered yet.
The human workaround is part of the risk model
One of the strongest points in the episode is that people react predictably when access control gets in the way of movement. They bypass it. They prop doors. They share credentials. They improvise access decisions because business still has to continue.
That matters because many organizations evaluate access control only in terms of system design, not human response. A backup process that requires too many calls, too many approvals, or too much ambiguity will be bypassed in the moment. James Rogers makes this especially clear when he notes that frontline staff need to be empowered and trained. If the approved process is too hard to execute, staff will choose the unofficial one.
In other words, resilience is not just whether the technology can keep working. It is whether the people around the technology know how to respond without making the situation worse.
Mechanical keys are useful and dangerous at the same time
The episode spends meaningful time on a common point of friction: should organizations rely on mechanical overrides during outages? Michael Harrington argues that mechanical keys solve one problem while creating another. Once a key is handed out, auditability can disappear. Teams may no longer know who used it, when it was used, or whether it was duplicated or left accessible in the wrong place.
James Rogers does not dismiss that concern, but he argues for a controlled mechanical option during emergencies. In a power outage, immediate safety and continuity may require something simpler than a digital workaround. The important distinction is governance. He recommends dual custody, signed logs, or sealed key boxes with controlled manager access.
That compromise is worth noting because it reflects a mature resilience posture. Mechanical backup is not wrong. Uncontrolled mechanical backup is the problem. If a key becomes a convenience tool instead of an emergency tool, the organization has weakened its own controls.
Why local caching needs tiered rules
The conversation around local credential caching is one of the most practical parts of the episode. Local caching allows readers to keep operating when the central server is unavailable. That can be the difference between a smooth short outage and a building-wide bottleneck. But cached permissions come with risk. If a credential should be revoked, suspended, or changed, those updates may not reach the reader until it reconnects.
The panel is clear that not every space should be treated the same way. A general office area may tolerate several hours of cached permissions. A more sensitive area should not. Michael pushes back on one-size-fits-all rules and argues for shorter cache windows and stronger fallback controls in high-security zones.
That tiering mindset is broadly useful beyond access control. It reminds facility and security leaders that resilience settings should follow business risk, not convenience alone. The safest and most workable design is often one that deliberately varies by zone.
Power resilience should also be tiered
When the discussion moves to UPS and lock power strategy, the speakers recommend a tiered approach there too. For perimeter doors and critical controllers, James recommends UPS coverage in the four-to-eight-hour range or enough time to bridge to generator transfer. That is practical guidance because it matches the parts of the system that create the biggest operational impact if they fail.
For interior doors, the focus shifts to battery-backed locks, remote battery reporting, and scheduled replacement cadence. But the episode does not stop at saying batteries should be checked. It stresses that vendor defaults are not enough. If alerts do not create tickets in the correct queue, if no one owns response, or if staff assume the system will speak for itself, the warning may be missed until the lock fails under real use.
That is an important operational insight: monitoring is only useful if it connects to responsibility.
Start with a 20-minute tabletop, not a giant program
Many organizations delay resilience work because they assume it requires a large project. This episode pushes in the opposite direction. James Rogers offers a 20-minute tabletop exercise that teams can run with minimal disruption.
The scenario is simple: a primary controller loses cloud connectivity during morning ingress.
- Minute 0: facilities verifies UPS and power conditions.
- Minutes 1 to 5: security tests cached credential behavior.
- Minutes 6 to 10: IT checks carrier and backend health.
- Minutes 11 to 15: the team practices the approved temporary access method.
- Minutes 16 to 20: the team debriefs, logs decisions, and updates the SOP.
The panel also recommends a clear decision trigger: if the issue is not resolved in 30 minutes, escalate to building management and send a tenant notification. That communication piece matters because uncertainty drives bad behavior. A quick, credible update can prevent frustration from turning into door propping and uncontrolled workarounds.
A small spare kit can prevent a big headache
Another practical takeaway is the recommendation for a lean on-site spare kit. Michael Harrington suggests keeping an indexed set of essentials that includes spare readers, a controller module, standard batteries, and documented firmware images. He also emphasizes documenting vendor roles and service-level expectations so there is no confusion over who is responsible when hardware needs to be replaced.
This is a strong reminder that resilience depends on logistics as much as architecture. A building can have a well-designed system and still lose time if the replacement part is unavailable, the image is missing, or the response chain is unclear.
Use telemetry, then adjust maintenance cadence
The episode includes a useful moment of disagreement around battery replacement cadence. Quarterly checks are proposed as a starting point, but Michael challenges the idea that quarterly is always enough, especially in mixed-tenant environments where some devices see much heavier use. James responds with a better rule: start with a baseline cadence, monitor battery drain remotely for the first three months, and then adjust to monthly or staggered replacements if the data shows faster discharge.
That exchange captures the operating philosophy of the entire episode. Start simple, but do not stay simplistic. Use measurement to refine the plan.
The three actions that matter most this week
By the close, the episode leaves listeners with a short list of actions that can improve resilience quickly:
- Document how critical doors fail and share that information across security, facilities, and IT.
- Verify that alerts from devices land in a monitored ticketing queue with a clear owner.
- Prepare a 30-minute escalation script and a tenant notification template.
- Run one 20-minute tabletop exercise with the people who will actually respond.
- Build a small spare kit and begin battery reporting so maintenance cadence follows real usage data.
The final caution is especially relevant for modern building environments: involve IT early. A physical backup that creates a network exposure is not really a win. Real resilience connects physical security and cyber discipline instead of treating them as separate conversations.
If your building or portfolio depends on electronic access systems, this episode is worth a listen because it turns a vague concern into a practical operating framework. The message is simple: schedule the tests, train the staff, organize the spares, and make failure response boring before the next Tuesday morning proves why that matters. Listen to the full episode here: https://builtwiredsecured.com/episodes/when-the-doors-stop-talking-access-control-resilience-without-drama