Skip to contentAnemo
EN
Contact

Database Migration Checklist: Planning, Testing, and Rollback

· 5 min read

A database migration succeeds or fails on three things decided before any data moves: what the acceptable data loss and downtime windows are, how the two systems are reconciled while both run, and what triggers a rollback. Research on migration projects puts the share that fail outright or significantly exceed budget at around 83%, and almost all of those failures are planning failures rather than technical ones.

This guide is a working checklist: assess, prepare, rehearse, cut over, verify, and keep the rollback real until you are certain.

Key takeaways

Assess before you plan

Record for the source system: total volume, growth rate, largest tables, character encoding, and every consumer that reads or writes it, including reports, integrations and the spreadsheet someone refreshes weekly. That last category is routinely missed and is where post-migration surprises come from.

Then profile the data honestly. Duplicates, orphaned records, invalid dates, fields used for something other than their name, and values that violate the constraints you plan to introduce. Every migration discovers that production data does not match the schema anyone believed in.

Decide explicitly what happens to bad records: cleaned before migration, migrated as-is, or quarantined for manual handling. Deciding this during the cutover window is how a four-hour migration becomes a twelve-hour one.

Set the two windows that govern everything

Acceptable downtime determines the approach. Hours available means a straightforward export and import. Minutes means replication with a short switchover. Zero means dual writes and a phased cutover, which is substantially more engineering.

Acceptable data loss, the recovery point objective, determines how the final sync works and what you promise the business. For most transactional systems the honest answer is zero, and that constrains the method.

Write both down and have them agreed by someone who can accept the consequences. These two numbers, more than any technical choice, decide the cost of the migration.

Prepare the target properly

Create the schema with the constraints, indexes and permissions you actually want, rather than mirroring the source's accumulated compromises. A migration is the cheapest opportunity you will get to fix a data model, and the most expensive one to skip.

Confirm character sets and collation explicitly. Encoding mismatches are the classic silent corruption: the migration reports success and Turkish characters are mangled in a subset of rows nobody checks until a customer complains.

Size the target for the load and leave headroom. Migration itself is often the heaviest write the target will ever take.

Rehearse at production scale

A rehearsal on a copy of production data is the single most valuable step, and the one most often cut for time.

It tells you the real duration, which is the number the cutover plan depends on and cannot be extrapolated from a sample. It surfaces the data problems that only exist at volume. It gives the team a script rather than an improvisation. And it lets you rehearse the rollback, which is otherwise theoretical.

Run it at least twice: once to find the problems, once to confirm the timings after fixing them.

Build reconciliation before you need it

Verification cannot be a spot check. Automate three layers:

Layer What it compares Catches
Row counts Per table, source versus target Wholesale omissions
Checksums or hashes Field-level across sampled or full sets Silent corruption and encoding issues
Business totals Order value, account balances, record counts by status Logic errors the technical checks pass

The third layer is the one that matters to the business and the one most often skipped. Matching row counts with wrong totals is a migration that technically succeeded.

Write the cutover as a runbook

A cutover plan is a timed sequence with named owners, not a summary. It should specify: when writes stop on the source, who confirms it, the final sync, the verification steps and who signs each off, when traffic switches, and the decision point.

Include the communications: who tells staff, what customers see during the window, and who is on call afterwards.

Schedule it when the business can absorb a problem, not the night before month-end close, and not immediately before the team's weekend.

Keep the rollback real

Rollback is only real if the source is still capable of taking traffic and you have rehearsed returning to it. Nearly one in five migrations needs that path.

Define the trigger in advance: what condition means abort, who is authorised to call it, and by when the decision must be made. Under pressure at 3am, a pre-agreed trigger is what prevents a slow slide into an unrecoverable position.

Keep the source system available and consistent for a defined period after cutover, a week is common. Do not repurpose its hardware the same day.

Verify before you decommission

Run the reconciliation, then watch the real system: error rates, slow queries, integration failures and support contacts. Some problems only appear when real users hit real paths, a report that returns nothing, an integration expecting an old field name, a nightly job that fails silently.

Agree the conditions for decommissioning the source, and a date. Systems kept "just in case" indefinitely become an unfunded maintenance and security liability, and a second place where data quietly diverges.

Should Database Migration Checklist: Planning, Testing, and Rollback become a modernisation programme rather than a series of fixes, Discuss Your Modernization Plan.

Frequently asked questions

What do we need to check before migrating our database?

Know what you are moving and what you are deliberately leaving, how the old fields map to the new ones, how long the business can be without the system, and how you would get back if the cutover fails. The mapping work is where unknown data quality problems surface, and it should happen weeks before the window.

How realistic does the rehearsal need to be?

Production scale, with production volume, on comparable infrastructure. A rehearsal on a tenth of the data proves the script runs and tells you nothing useful about how long it takes. The duration is what determines whether the window you negotiated is achievable.

When can we decommission the old database?

After reconciliation passes on the counts and on a sample of records, and after a full business cycle has run on the new system. A month end, a payroll run or a quarterly report will exercise code paths that daily use never touches, and that is exactly when you want the old data still available.

How we would work on this

Related services

Related reading

Building the product for what comes next

We would rather deliver one product that holds up than three that have to be rebuilt. That standard is the same on every project, whatever its size.

Ali Boran GazelCEO

Contact us