Skip to content

Reference architecture

Multi-region active-active database architecture

Keep serving reads and writes through a full regional outage by running a distributed SQL database active-active across three or more regions.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Near-zero recovery time (RTO) for a full regional outage
  • 02Recovery point objective of zero for committed writes (synchronous quorum replication)
  • 03Strict consistency for ledger transactions, with fast local reads
  • 04Automated, health-checked failover at the global routing layer

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

Multi-region active-active database architecture

Illustrative reference architecture

  1. 01

    Global traffic director

    Routes users to the nearest healthy region

    Cloudflare load balancing / Amazon Route 53 ARC

  2. 02

    Multi-region compute

    Identical application services in US, EU and Asia regions

    Kubernetes (EKS or GKE) in each region

  3. 03

    Distributed SQL database

    Synchronous, quorum-based replication across regions

    Google Cloud Spanner / Amazon Aurora DSQL

  4. 04

    Cross-region event bus

    Replicates domain events between regional Kafka clusters for downstream consumers

    Kafka MirrorMaker 2

Subsystem 01

Global traffic director

Routes users to the nearest healthy region

Typical stack

Cloudflare load balancing / Amazon Route 53 ARC

Subsystem 02

Multi-region compute

Identical application services in US, EU and Asia regions

Typical stack

Kubernetes (EKS or GKE) in each region

Subsystem 03

Distributed SQL database

Synchronous, quorum-based replication across regions

Typical stack

Google Cloud Spanner / Amazon Aurora DSQL

Subsystem 04

Cross-region event bus

Replicates domain events between regional Kafka clusters for downstream consumers

Typical stack

Kafka MirrorMaker 2

Data lifecycle

End-to-end data flow

  1. A request reaches the nearest edge (Anycast) and is routed to that region's compute cluster.

  2. Reads are served in-region: strong reads where correctness demands it, bounded-staleness reads where it doesn't.

  3. Writes commit through distributed consensus (Paxos or Raft) across regions.

  4. If a region fails its health checks, global routing stops sending traffic there; how fast depends on the health-check interval and DNS TTL.

  5. The surviving regions still hold a quorum, so they keep accepting writes with no manual database promotion.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Cross-region split brain

Mitigation

Use an odd number of voting regions (three or more), so only one side of a partition can hold a quorum.

Failure mode 02

High write latency across regions

Mitigation

Shard tenants by geography and place each shard's leader in its home region, so EU tenants' writes are led from Europe.

Failure mode 03

Intercontinental network partition

Mitigation

The side without a quorum stops accepting writes; it keeps serving stale reads and tells users clearly that changes are paused.

Questions

What teams ask about this design

Active-passive keeps a standby region idle until a failure. Active-active serves live traffic from every region at once, so there is no cold standby to promote.

Spanner's TrueTime API uses GPS receivers and atomic clocks to bound clock uncertainty, which lets it order transactions globally with external consistency.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.