Show Notes
Why Carrier Cutovers Create Building-Wide Risk
A carrier transition can look straightforward: schedule the change, move the circuit, confirm the internet is working, and close the ticket. In a commercial building, that approach can create a serious operational failure.
This episode examines a scenario in which a midsized office tower completed an after-hours cutover. Phones and guest Wi-Fi came online, but access control failed, the CCTV recorder lost streams, and a cloud-based elevator-monitoring feed dropped during morning traffic. The result was not a short maintenance event. It became two days of firefighting, tenant calls, facilities disruption, and credits.
The central lesson is simple: a circuit is not just an internet connection. It may support phones, physical security, building automation, elevator telemetry, point-of-sale systems, tenant cloud services, and other operational technology. A successful carrier cutover requires ownership, real validation, a defined rollback decision, and clear communication.
Start With a Single Cutover Owner
Carrier cutovers touch multiple teams: carrier personnel, IT, facilities, security, building operations, and sometimes tenant representatives. A committee can contribute expertise, but it should not be responsible for making the final coordination decisions.
- Name one cutover owner who coordinates the carrier, facilities, IT, and tenant liaison.
- Give that person responsibility for the runbook, schedule, escalation path, acceptance process, and rollback decision process.
- Ensure every participating party knows who has authority to move forward, pause work, or trigger rollback.
Clear ownership reduces delay when a problem appears during a maintenance window. It also prevents the common failure mode of each vendor assuming another party owns the issue.
Document Demarcation Points and Acceptance Signoffs
Demarcation clarity is one of the biggest risk reducers in a carrier transition. The team should document, in writing, where the carrier responsibility ends, where the building responsibility begins, and where tenant equipment becomes part of the path.
- Carrier demarcation: the point at which the carrier accepts responsibility for service delivery.
- Building demarcation: the handoff into building infrastructure and systems.
- Tenant equipment: the equipment and services that depend on the building and carrier path.
The episode recommends three distinct acceptance signoffs:
- Carrier acceptance at the carrier demarcation.
- Technical acceptance by IT or a systems engineer, based on connectivity and performance testing.
- Facilities or operations acceptance for building systems.
For tenant-critical services, provide a tenant acceptance window when requested. Include escalation contacts in the runbook before the cutover begins. Searching for a carrier contact or vendor escalation number during a failure wastes valuable time and extends tenant impact.
Respect Lead Times and Carrier Coordination Constraints
Carrier provisioning and testing often require weeks. When a schedule is compressed to meet a deadline, teams should acknowledge the trade-offs rather than assume the same outcome can be achieved with less preparation.
- Fewer available test cycles.
- A higher likelihood of needing to back out.
- More handoffs between vendors.
- Less time to identify routing, firewall, load, or performance issues before production use.
A tight schedule does not eliminate risk; it changes the amount of risk the organization accepts. Decision makers should understand that trade-off before approving the cutover window.
Use High-Value Tests That Reflect Real Operations
A basic ping test or proof that a browser can reach the internet is not enough. The cutover plan should begin with basic connectivity validation, then move quickly to application-level tests that exercise the real services people depend on.
- Validate connectivity to the internet and carrier-provided circuits.
- Place and receive VoIP calls across both PSTN and internal dial plans.
- Authenticate and exercise access-control readers.
- Stream CCTV recordings for 30 minutes to validate retention and latency.
- Run building automation system commands that control HVAC set points.
- Simulate a primary-circuit failure and confirm upstream failover works within the applicable SLA window.
These tests catch issues that a basic connectivity check can miss. The episode describes a rushed weekend cutover in which phones and internet passed basic tests, but a cloud-based access-control portal used a different outbound route and failed. Tenants could not badge in. The lesson: test actual user flows, including the routing and firewall requirements behind them.
Define Rollback Triggers Before the Window Opens
Rollback should never be an emotional decision made late at night after a series of improvised fixes. The runbook needs explicit conditions that tell the team when to stop troubleshooting and restore the prior state.
A practical rule from the episode is this: if any tenant-critical system fails a predefined acceptance test after two remediation attempts within the scheduled window, trigger rollback.
Examples include VoIP call quality dropping below MOS thresholds or access-control readers failing authentication when rebooting endpoints does not resolve the issue. The specific tests and thresholds should be written into the plan in advance, along with the precise steps and owners required to return service to the prior configuration.
Use an Observation Window After Cutover
Completion of the scheduled change is not the end of the work. The episode recommends an elevated-monitoring observation window, typically 24 to 72 hours, with a clearly identified rapid-response team.
- Update network diagrams, IP changes, circuit IDs, and contact lists.
- Monitor for performance, routing, latency, and service issues that may appear only under normal load.
- Send tenants a short status update when the change is complete.
- Send a second update at the end of the observation window.
Tenant communication is part of operational resilience. Clear, timely status updates reduce unnecessary escalation and help build trust when infrastructure work affects visible building services.
Validate in Parallel When Possible
One field example used a 48-hour parallel validation window. The new circuit operated in parallel while the old circuit remained active. During that period, the team found intermittent VoIP jitter under peak load and tuned QoS before final cutover.
Parallel validation can reveal load and latency problems that a single-point test does not expose. Where the environment and schedule allow it, this approach creates a more controlled path to the final transition.
Make Each Cutover Better Than the Last
Within a week of the change, conduct a lessons-learned review. Capture actual test results, unexpected issues, and the actions required to resolve them. Update the template and runbook so the next carrier transition benefits from what the team learned.
Long-term resilience also means building redundancy where feasible and validating failover as part of regular maintenance. Most importantly, store runbooks where operational teams can find them. A critical procedure buried in an email chain is not a dependable operational control.
Key Takeaways
- Name a single cutover owner and document demarcations, responsibilities, and signoffs.
- Test real user workflows, not only basic connectivity.
- Include failover validation and explicit rollback triggers in the runbook.
- Monitor for 24 to 72 hours after the transition and keep tenants informed.
- Capture lessons learned and improve the cutover template after every change.
Carrier Cutovers Need an Operations Playbook, Not a Connectivity Check
Carrier swaps and circuit transitions can appear deceptively simple. A new circuit is provisioned, a maintenance window is scheduled, the carrier completes its work, and the team confirms that internet access is available. For a commercial property or business environment, however, this narrow view misses the systems that depend on the connection.
Phones, guest Wi-Fi, access control, CCTV recording, point-of-sale systems, building automation systems, elevator telemetry, tenant cloud services, and other operational technology can all be affected by a circuit change. When a transition is planned only as a carrier event, a building can discover too late that a critical application takes a different route, depends on a firewall rule, or fails when placed under normal load.
A more reliable approach is a repeatable cutover playbook: one that establishes ownership, maps responsibilities, validates actual user workflows, defines rollback decisions before pressure builds, and monitors the environment after the work is complete.
Understand What Is Really Connected to the Circuit
The first question in a carrier transition should be: what breaks if this connection goes down or changes behavior? The answer is rarely limited to office internet access.
In the episode, an office tower scheduled an after-hours carrier cutover. Phones and guest Wi-Fi came up, which could have looked like a successful change. But access control failed, CCTV streams dropped, and a cloud-based elevator-monitoring feed was lost during morning traffic. What was expected to be a routine maintenance event became two days of troubleshooting, tenant calls, facilities response, and service credits.
The problem was not simply that a circuit changed. The problem was that the change was not validated against the full set of building services that depended on it. This is why planning must begin with an impact inventory. Identify phone systems, physical-security systems, building automation, tenant-facing services, and cloud services that may rely on the circuit or its routing path.
That inventory helps decision makers prioritize the systems that deserve explicit acceptance tests and rollback protection. It also prevents the assumption that a carrier will automatically understand every downstream application dependency in a building environment.
Assign One Cutover Owner
Carrier cutovers bring together several groups with different responsibilities: carrier teams, IT, facilities, physical security, building operations, and tenant contacts. Collaboration is necessary, but shared participation should not mean unclear ownership.
The playbook calls for a single cutover owner, not a committee. This person coordinates the carrier, facilities team, IT team, and tenant liaison. They manage the schedule, ensure that the runbook is ready, confirm that the right people are available, and keep the transition moving through validation or rollback.
A single owner does not eliminate technical accountability from the other teams. Instead, it creates a clear operational center. During a failure, there should be no question about who is coordinating vendor escalation, tracking remediation attempts, confirming acceptance results, or deciding whether a predefined rollback condition has been met.
For property and IT leaders, this is an important business control. Unclear ownership creates delay exactly when a maintenance window is shrinking and tenant impact is growing.
Write Down Demarcation Points and Signoffs
Demarcation points define where responsibility changes hands. They should be documented in writing for every transition: carrier demarcation, building demarcation, and tenant equipment.
This documentation matters because a circuit may be delivered successfully to a carrier demarcation while a problem remains in building infrastructure, routing, firewall policy, or tenant equipment. Without a documented handoff model, teams can spend hours debating responsibility rather than resolving the problem.
The episode recommends three stages of acceptance. First is carrier acceptance at its demarcation. Second is technical acceptance by IT or a systems engineer based on connectivity and performance testing. Third is facilities or operations acceptance for building systems. If tenant-critical services are involved, provide a tenant acceptance window when requested.
The runbook should also include current escalation contacts. A carrier number, facilities contact, security-system vendor contact, and relevant tenant liaison should be available before the work begins. An operational plan that requires teams to hunt through email during a service failure is incomplete.
Do Not Treat Compressed Lead Time as Normal
Carriers often need weeks for provisioning and testing. A tenant deadline or project schedule may create pressure to shorten that process, but compressing the window does not make the underlying dependencies disappear.
It reduces the number of test cycles, increases the chance of a backout, and creates more vendor handoffs under pressure. Leaders should view a compressed schedule as an explicit risk decision. If the project must proceed quickly, the plan should document the resulting limits and strengthen the controls that remain available, including clear acceptance tests, rapid escalation, and a well-defined rollback process.
The goal is not to avoid every schedule constraint. It is to avoid pretending that a constrained schedule provides the same confidence as a properly staged validation period.
Test Actual User Flows
A basic internet test can prove that some connectivity exists. It cannot prove that the services a building relies on will function correctly after a carrier transition.
The cutover sequence should begin with basic connectivity to the internet and carrier-provided circuits. It should then move into a small set of high-value, application-level tests. Those tests should be practical, repeatable, and connected to real operational outcomes.
- Place and receive VoIP calls across PSTN and internal dial plans.
- Authenticate and exercise access-control readers.
- Stream CCTV recordings for 30 minutes to validate retention and latency.
- Run building automation commands that control HVAC set points.
- Simulate a primary-circuit outage and confirm that upstream failover works within the SLA window.
A field example from the episode shows why this matters. A building rushed a weekend cutover to meet a tenant deadline. Phones and internet passed basic checks, but a cloud-based access-control portal used a different outbound route and failed. Tenants could not badge in. The lesson was direct: ping tests do not reveal application-level routing or firewall problems. Test the workflows that real people and systems will use.
Establish Rollback Triggers Before Troubleshooting Begins
Every transition needs a point at which the team stops trying incremental fixes and returns to the prior state. That point must be defined before the maintenance window begins.
The playbook offers a clear rule: if any tenant-critical system fails a predefined acceptance test after two remediation attempts within the scheduled window, trigger rollback. This protects the team from making exhausted, emotional decisions at 2 a.m. when the window is closing and the consequences of continued experimentation are increasing.
Examples might include VoIP call quality falling below MOS thresholds or access-control readers failing authentication after endpoint reboots do not resolve the problem. The exact criteria should be tailored to the environment, but they must be explicit. The runbook should specify the failed condition, the permitted remediation attempts, the authority to call rollback, and the steps required to restore the prior configuration.
Rollback is not an admission of failure. It is a planned resilience measure that protects building operations and gives teams time to diagnose the issue without leaving critical systems in an unstable state.
Use an Observation Window and Communicate Clearly
Carrier work may be complete at the end of the maintenance window, but the transition is not truly complete until the environment has operated reliably under normal conditions. The episode recommends an elevated-monitoring observation window of 24 to 72 hours, supported by a clearly identified rapid-response team.
During this period, update network diagrams, IP changes, circuit IDs, and contact lists. Monitor for performance, latency, routing, and service issues that may emerge only under regular tenant activity or peak load. Send a concise completion update to tenants after the change, then send another update when the observation window ends.
Clear communication has operational value. It helps tenants understand the status of services, reduces unnecessary escalation, and reinforces confidence that the building team is managing the change responsibly.
Parallel Validation Can Expose Problems Early
Where feasible, run the new circuit in parallel with the existing circuit before the final transition. In one example discussed in the episode, a 48-hour parallel validation period uncovered intermittent VoIP jitter under peak load. The team tuned QoS before final cutover.
This is a meaningful advantage over a one-time test. Parallel validation can reveal latency, load, and quality issues that are invisible during a brief maintenance-window check. It gives the team time to correct conditions before the old path is removed.
Turn Every Cutover Into a Better Runbook
Within a week of the change, conduct a lessons-learned review. Capture real test results, unexpected issues, remediation details, and any gaps in the contact list or escalation process. Update the cutover template so the next transition begins with better controls.
Over time, organizations can also improve resilience by building redundancy where practical and routinely validating failover. Just as importantly, runbooks need to be accessible. Procedures buried in email chains do not help an operations team during an outage.
A disciplined carrier-cutover process protects more than connectivity. It protects tenant experience, building operations, and the credibility of the teams responsible for critical technology. For a practical walkthrough of this approach, listen to this episode of Built, Wired & Secured and use the discussed checklist and sample runbook as a starting point for your next transition.