Subsystem 01
Global traffic director
Routes users to the nearest healthy region
Typical stack
Cloudflare load balancing / Amazon Route 53 ARC
Reference architecture
Keep serving reads and writes through a full regional outage by running a distributed SQL database active-active across three or more regions.
Design constraints
Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.
Component topology
Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.
Stack topology
Multi-region active-active database architecture
Illustrative reference architecture
Global traffic director
Routes users to the nearest healthy region
Cloudflare load balancing / Amazon Route 53 ARC
Multi-region compute
Identical application services in US, EU and Asia regions
Kubernetes (EKS or GKE) in each region
Distributed SQL database
Synchronous, quorum-based replication across regions
Google Cloud Spanner / Amazon Aurora DSQL
Cross-region event bus
Replicates domain events between regional Kafka clusters for downstream consumers
Kafka MirrorMaker 2
Subsystem 01
Routes users to the nearest healthy region
Typical stack
Cloudflare load balancing / Amazon Route 53 ARC
Subsystem 02
Identical application services in US, EU and Asia regions
Typical stack
Kubernetes (EKS or GKE) in each region
Subsystem 03
Synchronous, quorum-based replication across regions
Typical stack
Google Cloud Spanner / Amazon Aurora DSQL
Subsystem 04
Replicates domain events between regional Kafka clusters for downstream consumers
Typical stack
Kafka MirrorMaker 2
Data lifecycle
A request reaches the nearest edge (Anycast) and is routed to that region's compute cluster.
Reads are served in-region: strong reads where correctness demands it, bounded-staleness reads where it doesn't.
Writes commit through distributed consensus (Paxos or Raft) across regions.
If a region fails its health checks, global routing stops sending traffic there; how fast depends on the health-check interval and DNS TTL.
The surviving regions still hold a quorum, so they keep accepting writes with no manual database promotion.
Reliability and resilience
Failure mode 01
Mitigation
Use an odd number of voting regions (three or more), so only one side of a partition can hold a quorum.
Failure mode 02
Mitigation
Shard tenants by geography and place each shard's leader in its home region, so EU tenants' writes are led from Europe.
Failure mode 03
Mitigation
The side without a quorum stops accepting writes; it keeps serving stale reads and tells users clearly that changes are paused.
Questions
Active-passive keeps a standby region idle until a failure. Active-active serves live traffic from every region at once, so there is no cold standby to promote.
Spanner's TrueTime API uses GPS receivers and atomic clocks to bound clock uncertainty, which lets it order transactions globally with external consistency.
Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.