Show Notes
Why "Two ISPs" Still Fails in Real Life
This episode of Built, Wired, and Secured takes on a problem that shows up in commercial real estate and multi-tenant operations all the time: a building appears redundant on paper, but a single failure still knocks everything offline. The opening scenario makes the risk concrete. Access control alarms trigger, tenant VPNs fall over, and two different carrier portals both show outages. The building has two providers listed on the floor plan, yet operations are still down. The core lesson is simple: separate carrier names do not automatically mean real resilience.
Alex Morgan is joined by Michael Harrington and James Rogers to break down the difference between logical diversity and physical diversity. Logical diversity means different services or providers. Physical diversity means those services do not share the same riser, conduit, sleeve, entry path, or meet-me room. If they do, a single cable cut, maintenance error, or pathway issue can take out both links at once. That is the "paperwork illusion" the guests warn against.
What Real Carrier Diversity Requires
The conversation quickly moves from theory to execution. True redundancy starts with disciplined mapping, verification, and ownership. The team explains that organizations need to know exactly where each cable runs, whether separate carriers share the same physical pathways, and who owns each demarcation point. The advice is direct: trust the process, but verify in person.
- Map physical routes, not just service names
- Identify shared sleeves, conduits, risers, and meet-me rooms
- Document every demarc and handoff responsibility
- Verify carrier claims on site before cutover
The guests also stress that not every system needs the same level of protection. The practical approach is to prioritize the most critical services first. Life safety, building automation, and high-impact tenant systems should drive the first investment decisions. In other words, design diversity around business impact, not around a generic idea of redundancy.
The Meet-Me Room and Demarc Problem
One of the strongest operational warnings in the episode centers on the meet-me room. Contracts may promise diverse handoffs, but those handoffs can still terminate only a few feet apart in the same crowded closet. That creates a hidden single point of failure. A good procurement spec, the guests explain, should require annotated meet-me room details, route diagrams, and explicit inspection rights. Then someone has to physically confirm that the installed reality matches the paperwork.
This is also where internal ownership often breaks down. Facilities may assume IT is validating carrier paths. IT may assume the carrier has already done the verification. The carrier may only be confirming its own side of the handoff. The answer offered in the episode is clear accountability.
- Facilities typically owns pathways, risers, and telecom spaces
- IT owns service validation and failover behavior
- Carrier contracts should require route maps and demarc documentation
- One named coordinator should own inspections and test execution
That single coordinator role matters because it prevents finger-pointing during projects and outages.
Why Testing Failover Changes Everything
The episode makes a strong case that redundancy is unproven until it has been tested. Some teams avoid failover tests because they fear disruption. Others test once during implementation and never revisit it. The guests recommend staged, scheduled failovers that start small and expand only after each step is validated.
The testing guidance is practical and repeatable:
- Start with non-critical systems or a limited scope such as one floor
- Use agreed maintenance windows
- Notify tenants in advance
- Have rollback steps ready before the test begins
- Keep carriers on the line during the event
- Use monitoring to confirm traffic truly shifts to the backup path
The goal is not to create drama. It is to create confidence. A five-minute controlled test can reveal issues that would be far more painful during a real outage.
The Cross-System Surprise Most Teams Miss
One of the most useful examples in the discussion involves a failover that looked healthy from the carrier side but still broke services. The backup circuit was live, yet a tenant firewall only allowed the primary carrier's IP ranges. When the switchover happened, applications failed. What looked like a carrier issue was actually a configuration and coordination issue spanning tenant systems, building operations, and network policy.
The fix was small but important: add secondary carrier IP ranges to tenant onboarding and building runbooks. That example reinforces a bigger point from the episode: resilience is not just about circuits. It is about the surrounding operational details that determine whether failover actually works under pressure.
Single Carrier vs. True Diversity
The conversation also addresses a common owner question: why not buy one excellent carrier with a strong SLA and skip the extra complexity? The answer is that SLAs improve repair expectations, but they do not provide instant continuity. A strong contract does not stop a conduit cut, a maintenance mistake, or physical damage. If the business requires systems to stay up in the moment, a physically separate backup path matters more than a promise of faster restoration.
That does not mean every building must buy maximum redundancy. It means the decision should be explicit, documented, and tied to acceptable downtime for critical systems. If leadership chooses a single-carrier design, that choice belongs in the risk register with signoff from the right stakeholders.
The One-Page Checklist to Use Next
The episode closes with a short, actionable checklist listeners can use during procurement, retrofits, and outage drills. The emphasis is on simple, repeatable operating discipline rather than complicated redesigns.
- Annotate physical routes, meet-me rooms, and demarc locations
- Name who owns pathways and who owns service testing
- Create a one-page runbook with carrier contacts, cable IDs, and failover steps
- Require physical route confirmations from carriers
- Conduct on-site verification before cutover
- Run staged failovers at least quarterly
- Define acceptable downtime for critical systems before making design choices
The closing takeaway is one that applies across building operations and infrastructure management: small repeatable checks beat big infrequent overhauls. Real carrier diversity is not a line item. It is a capability that has to be mapped, verified, tested, and maintained.
Carrier Diversity Only Works When It Is Real
Many buildings believe they are protected from connectivity outages because they can point to two internet providers on a floor plan or in a contract folder. It sounds reassuring. Two carriers should mean redundancy. But as this episode of Built, Wired, and Secured makes clear, that assumption often falls apart the moment a real incident happens.
The scenario that opens the episode is instantly recognizable to anyone responsible for building operations, tenant experience, or IT continuity. Midafternoon, alarms start hitting. Tenant VPNs drop. Carrier portals show multiple outages. On paper, the site has two ISPs. In practice, everything is down. The central problem is not that redundancy was ignored. It is that redundancy was defined too loosely to be useful.
The discussion with Michael Harrington and James Rogers separates a checkbox mindset from an operational one. The real issue is the difference between logical diversity and physical diversity. Logical diversity means there are two services, or even two different providers. Physical diversity means those services do not depend on the same vulnerable path. If two carriers use the same riser, the same conduit, the same sleeve, or terminate in the same meet-me room, a single event can still take both out.
That is why the episode calls this the paperwork illusion. Different carrier names can create false confidence when the underlying physical path is still shared.
Redundancy Should Be Designed Around Failure Impact
One of the best parts of the conversation is that it does not treat diversity as an all-or-nothing engineering exercise. The guests acknowledge the budget reality immediately. Full separation at every layer can cost more. Separate entrances, separate risers, and truly distinct physical routes all add complexity. So where should organizations focus?
The answer is to start with impact. Identify the systems where failure hurts most, then protect those first. Life safety systems, building automation, access control, and critical tenant services should drive the design conversation. Instead of asking, "How much redundancy can we afford?" a better question is, "What breaks if this path fails, and what is the business consequence?"
That shift matters. It moves the decision out of abstract technical preference and into operational risk management. For commercial real estate owners and facilities teams, that means thinking beyond bandwidth and contract language. It means understanding which outages interrupt tenant operations, trigger support escalations, affect safety, or damage trust.
The Meet-Me Room Is Often the Hidden Single Point of Failure
The episode is especially sharp on one issue that is easy to miss during procurement: the meet-me room. A contract can say diverse handoffs. A provider can present separate services. But if both services land a few feet apart in the same crowded telecom closet, the design may still share the same operational risk.
That is why route confirmation cannot stay at the level of diagrams and vendor claims. The guests argue for annotated meet-me room details, route diagrams, and inspection rights as part of the procurement and deployment process. More importantly, they recommend physically verifying the installation before cutover.
This is where many organizations run into a coordination gap. Facilities may control pathways, risers, and closets. IT may own service validation and network failover logic. Carriers own their portions of the route and handoff. Each group assumes someone else is checking the whole picture. The result is incomplete validation and a high chance of discovering shared dependencies only after an outage.
The practical fix is simple: assign one named coordinator. That person does not have to perform every technical task personally, but they do need clear accountability for inspections, route confirmation, and testing. Without that role, follow-through gives way to finger-pointing.
Testing Is What Turns Redundancy Into Capability
The episode repeatedly returns to a critical truth: failover that has not been tested is only an assumption. This is where organizations often hesitate. Some teams fear that testing backup paths could disrupt tenants. Others test once during implementation, document the result, and never revisit it. Neither approach builds confidence.
The recommended model is staged failover testing. Start with a narrow scope. Use non-critical systems or a single floor. Schedule an agreed maintenance window. Notify affected stakeholders. Keep rollback steps ready. Keep the carriers on the line. Then use monitoring to verify that traffic and services actually shift to the intended backup path.
This matters because a failover event tests more than circuits. It tests routing behavior, device logic, tenant dependencies, communication workflows, and operational readiness. If any of those pieces are weak, the backup path may exist technically while still failing in practice.
The guests give a strong real-world example. In one test, the secondary carrier was up, but a tenant firewall allowed only the primary carrier's IP ranges. When traffic shifted, services failed. The carrier was not the real problem. The configuration around the carrier was. That is exactly the kind of issue a short staged test can uncover early and cheaply.
The corrective action in that case was straightforward: add secondary carrier IP ranges to tenant onboarding and to the building runbook. The lesson is broader than firewall rules. Successful resilience depends on documenting the surrounding dependencies that determine whether failover will actually succeed during an incident.
Runbooks Need to Be Short Enough to Use
Another practical point from the episode is the value of a concise runbook. During an outage or a drill, no one wants to sort through a long binder of generalized documentation. The recommended format is a one-page operational guide containing the essentials.
- Demarc locations
- Cable IDs
- Carrier escalation contacts
- Scheduled maintenance windows
- A clear failover checklist
The emphasis on brevity is intentional. A runbook that is easy to carry into a drill is more likely to be used correctly under pressure. In operational environments, simple and reliable beats comprehensive but ignored.
SLAs Are Not the Same as Continuity
The episode also addresses a common budget question: if a premium carrier offers strong service levels, why not depend on that instead of paying for a diverse backup? The answer is one many operators already know from experience. Service-level agreements help define restoration expectations, but they do not keep a system online at the moment of failure.
A backhoe does not care about an SLA. Neither does a rodent in a conduit or a maintenance mistake in a shared pathway. If the business needs continuous operation for critical services, the better solution is physical separation or a proven backup path. If leadership decides the cost of that protection is not justified, the episode recommends documenting that decision explicitly in the risk register and obtaining the right signoffs.
That framing is important because it turns a vague compromise into an accountable business decision. It also forces everyone involved to define acceptable downtime for the systems that matter most.
What Teams Should Do Next
The closing checklist from the episode is useful because it is realistic. It does not assume a blank-slate construction project or an unlimited budget. It gives operators a practical sequence they can apply to the next procurement, retrofit, or outage drill.
- Map physical routes, meet-me rooms, and demarc locations
- Confirm who owns pathways and who owns service testing
- Require route confirmation from carriers
- Verify the physical reality on site before cutover
- Create a one-page runbook with contacts, IDs, and failover steps
- Run staged failovers at least quarterly
- Define acceptable downtime so decisions reflect business impact
The larger message is that resilient connectivity is operational, not cosmetic. Carrier diversity should not exist only in drawings, procurement language, or monthly invoices. It should show up in mapped routes, verified handoffs, tested failovers, clear ownership, and usable documentation.
If your building or portfolio relies on connectivity to support tenant operations, safety systems, and day-to-day business continuity, this episode offers a clear standard to work from. The next time someone says, "We already have two providers," the right follow-up is no longer whether there are two names on paper. It is whether those two paths can survive a real failure independently. If you want a better framework for that conversation, this episode is worth the listen.