Skip to content

Reference architecture

Zero-downtime database migration and replication

A method for moving a multi-terabyte production database to a new engine or host while it keeps serving traffic, with a rollback path at every stage.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Reads and writes keep working throughout; the cutover pause is measured in seconds
  • 02Rollback stays possible at every stage until the old database is retired
  • 03Full parity checks of data consistency between the source and target engines
  • 04Minimal extra load on the production source database

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

Zero-downtime database migration and replication

Illustrative reference architecture

  1. 01

    Change data capture (CDC)

    Streams committed changes from the source database's transaction log

    Debezium + Kafka Connect

  2. 02

    Replication buffer

    Holds ordered change events while the historical backfill runs

    Apache Kafka

  3. 03

    Target database engine

    Receives the replicated change stream

    Amazon Aurora PostgreSQL / PlanetScale

  4. 04

    Shadow verification service

    Compares source and target query results in real time

    Go shadow proxy

Subsystem 01

Change data capture (CDC)

Streams committed changes from the source database's transaction log

Typical stack

Debezium + Kafka Connect

Subsystem 02

Replication buffer

Holds ordered change events while the historical backfill runs

Typical stack

Apache Kafka

Subsystem 03

Target database engine

Receives the replicated change stream

Typical stack

Amazon Aurora PostgreSQL / PlanetScale

Subsystem 04

Shadow verification service

Compares source and target query results in real time

Typical stack

Go shadow proxy

Data lifecycle

End-to-end data flow

  1. An initial snapshot of the source database is restored to the target.

  2. Debezium streams inserts, updates and deletes from the transaction log (WAL or binlog) into Kafka.

  3. A Kafka sink connector applies the changes to the target until replication lag stays consistently near zero.

  4. Application reads are duplicated in shadow mode to confirm identical results and acceptable performance.

  5. Writes switch to the target, and replication is reversed so the old database stays current for rollback.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Replication lag spikes under heavy write bursts

Mitigation

Add Kafka partitions and tune the target's batch write sizes to keep replication lag under a second.

Failure mode 02

Primary key sequence desynchronization

Mitigation

Pre-cutover scripts verify and reset PostgreSQL sequences to MAX(id) on the target.

Failure mode 03

Foreign key lock contention during backfill

Mitigation

Drop foreign keys for the bulk load, re-add them as NOT VALID, and run VALIDATE CONSTRAINT before cutover.

Questions

What teams ask about this design

Replication runs in reverse from the new database to the old one, so the old database stays current and you can switch back with the writes made since cutover already applied.

Usually only a little, because CDC reads the transaction log instead of querying tables. The costs to watch are WAL retention behind a slow replication slot and the initial snapshot; we measure both on a staging copy before touching production.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.