Migration · Runbook
Cutover runbook
One migration wave, from the two-week freeze to the retrospective. The times below assume a weekend window and a wave of 10–30 racks; larger waves get split rather than stretched. Two things make this a runbook rather than a plan: every step has one named owner, and the go/no-go criteria are decided weeks before the night nobody is thinking clearly.
| T | When | Action | Owner | Detail |
|---|---|---|---|---|
| T-14d | Two weeks out | Change approved, wave locked | Migration lead | Change request approved, wave contents frozen. Anything not on the list at T-14 does not move in this wave — late additions are how a clean wave becomes an incident. |
| T-7d | One week out | Target racked, cabled, powered, reachable | Field engineering | Replacement or receiving hardware is in place at the target, on the right circuits, with out-of-band access proven from outside the corporate network. |
| T-72h | Tuesday | Dry run of the validation script | Application owners | Run the post-move checks against the current production estate. If a check fails now, it is not a migration failure at 2 a.m. on Saturday — it is a broken check. |
| T-48h | Wednesday | DNS TTLs lowered to 300s | Network engineering | Lowered far enough ahead that every resolver has picked up the short TTL before the cutover. |
| T-24h | Thursday | Full backup taken and restore-tested | Backup team | Verified restorable, not merely completed. A backup nobody has restored is a belief, not a control. |
| T-4h | Friday 17:00 | Change freeze begins, bridge opens | Migration lead | All unrelated changes stopped across the estate. Bridge call opened with infrastructure, network, application owners and the on-site team. |
| T-0 | Friday 21:00 | Go / no-go #1 | Migration lead + business sponsor | Criteria checked out loud: backup verified, target reachable, full team present, no active P1 elsewhere. Any single no is a no. Deferring costs one weekend; proceeding on a shaky no costs the quarter. |
| T+0:15 | Friday 21:15 | Graceful shutdown, final delta sync | Application owners | Applications stopped in dependency order, last replication delta flushed, source systems marked read-only so nothing writes to a site that is about to be empty. |
| T+1:00 | Friday 22:00 | Decommission from source, load, transport | Field engineering | Rails and cables labelled as they come out. Photographs of the front and rear of every cabinet before anything is unplugged — the fastest rollback aid there is. |
| T+4:00 | Saturday 01:00 | Receive, rack, cable, power on | Field engineering | Power on in staged groups rather than all at once, so a tripped breaker identifies itself instead of taking the row down. |
| T+7:00 | Saturday 04:00 | Go / no-go #2 — the rollback clock | Migration lead | The last point at which a rollback still fits inside the window. If the estate is not powered and reachable by this fixed clock time, roll back — the decision was made at T-14, not now. |
| T+8:00 | Saturday 05:00 | Network cutover, DNS repointed | Network engineering | Routing moved, firewall policy activated, DNS records repointed. Monitoring should light up green from the new site before anyone is told it worked. |
| T+10:00 | Saturday 07:00 | Automated validation pass | Application owners | The same script from the dry run, now against the target. Pass or fail, not opinion. |
| T+14:00 | Saturday 11:00 | Business validation and sign-off | Business sponsor | Named users exercise real transactions. Sign-off is written, per application, and belongs to the owner rather than the migration team. |
| T+24:00 | Sunday 21:00 | Freeze lifted, hypercare begins | Migration lead | Freeze released, elevated support for five business days, and the source environment left intact until hypercare closes. |
| T+7d | Following Friday | Retrospective, runbook updated | Migration lead | Corrections written into the runbook while the detail is fresh. Every subsequent wave inherits them. |
No-go criteria
Agreed at change approval, read aloud at each go/no-go. Any single one is a stop. The point of writing them down early is that at 4 a.m. the argument is already settled.
- Backup not verified restorable
- Target site not reachable over out-of-band
- A named owner missing from the bridge
- An active P1 anywhere in the estate
- Carrier circuit not confirmed live
- Rollback path untested since the last change
The rollback clock
A rollback is only real if it fits in the remaining window. Time it during the pilot: how long to re-rack, re-cable, power on and re-point DNS at the source site. Subtract that from the end of the window and you have a fixed clock time — the second go/no-go. Reaching it without a powered, reachable target means rolling back, regardless of how close the team feels to done. Teams that skip this step do not avoid rollbacks; they discover at 7 a.m. that they no longer have the option.
What to bring on the night
- Rack elevations and cabling diagrams, printed — the network you need to consult may be the one you are moving
- Photographs of every cabinet front and rear, taken before anything was unplugged
- Spare optics, patch cables, rails, PDUs and at least one spare of anything single-sourced
- Out-of-band access credentials tested from a phone tether, not the office network
- The escalation tree with mobile numbers, including the operator's remote hands desk
- The validation script, already proven against production in the dry run
Before the runbook, the plan
A cutover only goes well when discovery went well. Work the migration checklist first, model the programme with the cost calculator, and shortlist the target from the facility catalog.
Quotes for the target site
Tell us the wave size and the market — we return facilities that fit and benchmark pricing.

