▮▮Coloprice

Migration · Runbook

Cutover runbook

One migration wave, from the two-week freeze to the retrospective. The times below assume a weekend window and a wave of 10–30 racks; larger waves get split rather than stretched. Two things make this a runbook rather than a plan: every step has one named owner, and the go/no-go criteria are decided weeks before the night nobody is thinking clearly.

TWhenActionOwnerDetail
T-14dTwo weeks outChange approved, wave lockedMigration leadChange request approved, wave contents frozen. Anything not on the list at T-14 does not move in this wave — late additions are how a clean wave becomes an incident.
T-7dOne week outTarget racked, cabled, powered, reachableField engineeringReplacement or receiving hardware is in place at the target, on the right circuits, with out-of-band access proven from outside the corporate network.
T-72hTuesdayDry run of the validation scriptApplication ownersRun the post-move checks against the current production estate. If a check fails now, it is not a migration failure at 2 a.m. on Saturday — it is a broken check.
T-48hWednesdayDNS TTLs lowered to 300sNetwork engineeringLowered far enough ahead that every resolver has picked up the short TTL before the cutover.
T-24hThursdayFull backup taken and restore-testedBackup teamVerified restorable, not merely completed. A backup nobody has restored is a belief, not a control.
T-4hFriday 17:00Change freeze begins, bridge opensMigration leadAll unrelated changes stopped across the estate. Bridge call opened with infrastructure, network, application owners and the on-site team.
T-0Friday 21:00Go / no-go #1Migration lead + business sponsorCriteria checked out loud: backup verified, target reachable, full team present, no active P1 elsewhere. Any single no is a no. Deferring costs one weekend; proceeding on a shaky no costs the quarter.
T+0:15Friday 21:15Graceful shutdown, final delta syncApplication ownersApplications stopped in dependency order, last replication delta flushed, source systems marked read-only so nothing writes to a site that is about to be empty.
T+1:00Friday 22:00Decommission from source, load, transportField engineeringRails and cables labelled as they come out. Photographs of the front and rear of every cabinet before anything is unplugged — the fastest rollback aid there is.
T+4:00Saturday 01:00Receive, rack, cable, power onField engineeringPower on in staged groups rather than all at once, so a tripped breaker identifies itself instead of taking the row down.
T+7:00Saturday 04:00Go / no-go #2 — the rollback clockMigration leadThe last point at which a rollback still fits inside the window. If the estate is not powered and reachable by this fixed clock time, roll back — the decision was made at T-14, not now.
T+8:00Saturday 05:00Network cutover, DNS repointedNetwork engineeringRouting moved, firewall policy activated, DNS records repointed. Monitoring should light up green from the new site before anyone is told it worked.
T+10:00Saturday 07:00Automated validation passApplication ownersThe same script from the dry run, now against the target. Pass or fail, not opinion.
T+14:00Saturday 11:00Business validation and sign-offBusiness sponsorNamed users exercise real transactions. Sign-off is written, per application, and belongs to the owner rather than the migration team.
T+24:00Sunday 21:00Freeze lifted, hypercare beginsMigration leadFreeze released, elevated support for five business days, and the source environment left intact until hypercare closes.
T+7dFollowing FridayRetrospective, runbook updatedMigration leadCorrections written into the runbook while the detail is fresh. Every subsequent wave inherits them.

No-go criteria

Agreed at change approval, read aloud at each go/no-go. Any single one is a stop. The point of writing them down early is that at 4 a.m. the argument is already settled.

  • Backup not verified restorable
  • Target site not reachable over out-of-band
  • A named owner missing from the bridge
  • An active P1 anywhere in the estate
  • Carrier circuit not confirmed live
  • Rollback path untested since the last change

The rollback clock

A rollback is only real if it fits in the remaining window. Time it during the pilot: how long to re-rack, re-cable, power on and re-point DNS at the source site. Subtract that from the end of the window and you have a fixed clock time — the second go/no-go. Reaching it without a powered, reachable target means rolling back, regardless of how close the team feels to done. Teams that skip this step do not avoid rollbacks; they discover at 7 a.m. that they no longer have the option.

What to bring on the night

Before the runbook, the plan

A cutover only goes well when discovery went well. Work the migration checklist first, model the programme with the cost calculator, and shortlist the target from the facility catalog.

Quotes for the target site

Tell us the wave size and the market — we return facilities that fit and benchmark pricing.

We reply within one business day. No spam, no reselling your contacts.