Skip to content

Reference architecture

High-performance feature flagging and experimentation engine

Feature flags and A/B experiments evaluated in application memory, so a flag check never waits on the network and keeps working when the flag service is down.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Sub-millisecond in-memory flag evaluation with no network round trip
  • 02Rule changes reach every SDK within seconds of a toggle
  • 03Deterministic MurmurHash3 bucketing, so each user sees a consistent variant
  • 04Applications keep running on last-known rules if the flag server is unreachable

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

High-performance feature flagging and experimentation engine

Illustrative reference architecture

  1. 01

    Management console

    Create flags, set rollout percentages and target user segments

    Next.js + PostgreSQL

  2. 02

    Streaming distribution

    Pushes rule updates to application servers over SSE

    Server-Sent Events (SSE) / Redis Pub/Sub

  3. 03

    In-memory client SDK

    Evaluates targeting rules locally in application memory

    Go / TypeScript / Python SDKs

  4. 04

    Experimentation analytics

    Bayesian significance on conversion metrics

    ClickHouse + dbt

Subsystem 01

Management console

Create flags, set rollout percentages and target user segments

Typical stack

Next.js + PostgreSQL

Subsystem 02

Streaming distribution

Pushes rule updates to application servers over SSE

Typical stack

Server-Sent Events (SSE) / Redis Pub/Sub

Subsystem 03

In-memory client SDK

Evaluates targeting rules locally in application memory

Typical stack

Go / TypeScript / Python SDKs

Subsystem 04

Experimentation analytics

Bayesian significance on conversion metrics

Typical stack

ClickHouse + dbt

Data lifecycle

End-to-end data flow

  1. An engineer rolls a feature out to 20% of users in the management console.

  2. The change is written to PostgreSQL and broadcast over Redis Pub/Sub to the SSE servers.

  3. SDKs hold persistent SSE connections and update their in-memory rules as changes arrive.

  4. On each request, the SDK computes MurmurHash3(user_id + flag_key) % 100 locally, in microseconds.

  5. Evaluation events are buffered and flushed asynchronously to ClickHouse for significance tracking.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Management server outage

Mitigation

SDKs cache rule sets on local disk and keep evaluating flags from the last-known rules.

Failure mode 02

Network partition during a toggle

Mitigation

Automatic reconnection with exponential backoff, falling back to the last-known-good rules.

Failure mode 03

Stale flag code piling up

Mitigation

An AST-based linter in CI reports flags that are fully rolled out or retired.

Questions

What teams ask about this design

A network call per flag check adds a round trip, often tens of milliseconds, to every request that checks a flag, and makes the flag service a dependency of every page. Local evaluation takes microseconds and keeps working when that service is down.

MurmurHash3 is deterministic: the same user ID and flag key always produce the same hash, so a user lands in the same bucket on every request as long as the flag key and the variant split stay the same.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.