Skip to content

Major database migrations postponed for fear of long outages

Near-zero-downtime database migration playbook

Move a production database without a long maintenance window: stream changes with CDC, verify parity with shadow reads, and cut over in a short, rehearsed window with reverse replication ready for rollback.

Symptoms

Signs your platform has this problem

If several of these sound familiar, the plan below is where we would start.

01

Fear of long outages

Postponing database migrations because a dump and restore would take many hours.

02

Risk of inconsistent data

No automated way to check that millions of rows in the new database match the old one.

03

No way back

No rollback plan if performance problems appear after cutover.

Remediation plan

How we fix it, step by step

Each phase ends with a measurement, so you can see what changed before the next one starts.

01

Target sizing and schema translation

Translating schemas, foreign keys and indexes to the target PostgreSQL or Aurora, for example with pgloader.

02

Change data capture

Streaming transaction log changes into Kafka and the target database with Debezium.

03

Shadow read verification

Replaying production reads against the target to compare plans, latency and results.

04

Cutover and reverse replication

Switching writes in a short, rehearsed window and starting reverse replication so you can roll back.

Technical checklist

Remediation checklist

What we check before a change goes to production:

  • Configure Debezium CDC on the source transaction log (WAL or binlog)
  • Run row-count and checksum parity scripts across all tables
  • Verify sequences are ahead of MAX(id) before cutover
  • Set up reverse replication to the old database for rollback

What we measure

We take a baseline first and report the same measurements after each change, from your own monitoring — evidence, not promised results.

Write freeze
Measured in rehearsal, then at the real cutover
Row parity
Row counts and checksums per table before switching
Rollback time
Time to switch back, measured in a rehearsal

Related service

IT consulting

Senior counsel for decisions too expensive to get wrong — architecture reviews, technical due diligence, and delivery audits in writing, within weeks.

Explore IT consulting

Questions

Questions about this remediation

Reverse replication streams changes from the new database back to the old one, so switching back doesn't lose writes made after the cutover. We rehearse the rollback as well as the cutover.

Only a little: log-based CDC reads the WAL or binlog instead of querying your tables. It does hold a replication slot and add some I/O, so we monitor lag and disk on the source throughout.

Want an engineer to look at this with you?

Send us the symptoms and any metrics you have. We'll reply within one business day, set up a call and agree what to measure before anything changes.