Subsystem 01
Change data capture (CDC)
Streams committed changes from the source database's transaction log
Typical stack
Debezium + Kafka Connect
Reference architecture
A method for moving a multi-terabyte production database to a new engine or host while it keeps serving traffic, with a rollback path at every stage.
Design constraints
Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.
Component topology
Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.
Stack topology
Zero-downtime database migration and replication
Illustrative reference architecture
Change data capture (CDC)
Streams committed changes from the source database's transaction log
Debezium + Kafka Connect
Replication buffer
Holds ordered change events while the historical backfill runs
Apache Kafka
Target database engine
Receives the replicated change stream
Amazon Aurora PostgreSQL / PlanetScale
Shadow verification service
Compares source and target query results in real time
Go shadow proxy
Subsystem 01
Streams committed changes from the source database's transaction log
Typical stack
Debezium + Kafka Connect
Subsystem 02
Holds ordered change events while the historical backfill runs
Typical stack
Apache Kafka
Subsystem 03
Receives the replicated change stream
Typical stack
Amazon Aurora PostgreSQL / PlanetScale
Subsystem 04
Compares source and target query results in real time
Typical stack
Go shadow proxy
Data lifecycle
An initial snapshot of the source database is restored to the target.
Debezium streams inserts, updates and deletes from the transaction log (WAL or binlog) into Kafka.
A Kafka sink connector applies the changes to the target until replication lag stays consistently near zero.
Application reads are duplicated in shadow mode to confirm identical results and acceptable performance.
Writes switch to the target, and replication is reversed so the old database stays current for rollback.
Reliability and resilience
Failure mode 01
Mitigation
Add Kafka partitions and tune the target's batch write sizes to keep replication lag under a second.
Failure mode 02
Mitigation
Pre-cutover scripts verify and reset PostgreSQL sequences to MAX(id) on the target.
Failure mode 03
Mitigation
Drop foreign keys for the bulk load, re-add them as NOT VALID, and run VALIDATE CONSTRAINT before cutover.
Questions
Replication runs in reverse from the new database to the old one, so the old database stays current and you can switch back with the writes made since cutover already applied.
Usually only a little, because CDC reads the transaction log instead of querying tables. The costs to watch are WAL retention behind a slow replication slot and the initial snapshot; we measure both on a staging copy before touching production.
Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.