Skip to content

Reference architecture

High-reliability enterprise webhook dispatch system

Customer-facing webhooks with signed payloads, retries with backoff and jitter, and delivery logs developers can inspect and replay themselves.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01At-least-once delivery for every platform state event (consumers deduplicate by event ID)
  • 02Slow customer endpoints can't tie up internal systems (slow-loris timeouts)
  • 03HMAC-SHA256 payload signatures that customers can verify
  • 04Self-serve portal with delivery logs and manual redelivery

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

High-reliability enterprise webhook dispatch system

Illustrative reference architecture

  1. 01

    Event publisher

    Captures domain events and queues webhook deliveries

    FastAPI / Node.js + Redis

  2. 02

    Distributed dispatch workers

    Asynchronous HTTP delivery with strict concurrency limits

    Go worker pool + Temporal

  3. 03

    Webhook delivery log store

    Request headers, response bodies and latency for every attempt

    PostgreSQL / ClickHouse

  4. 04

    Developer webhook portal

    Endpoint testing, payload inspection and retry history

    Next.js App Router

Subsystem 01

Event publisher

Captures domain events and queues webhook deliveries

Typical stack

FastAPI / Node.js + Redis

Subsystem 02

Distributed dispatch workers

Asynchronous HTTP delivery with strict concurrency limits

Typical stack

Go worker pool + Temporal

Subsystem 03

Webhook delivery log store

Request headers, response bodies and latency for every attempt

Typical stack

PostgreSQL / ClickHouse

Subsystem 04

Developer webhook portal

Endpoint testing, payload inspection and retry history

Typical stack

Next.js App Router

Data lifecycle

End-to-end data flow

  1. The application publishes a domain event (such as payment.succeeded) to an internal Redis queue.

  2. A dispatch worker looks up the customer's registered endpoints and signing secrets.

  3. The worker computes an HMAC-SHA256 signature and attaches a timestamped signature header, in the style Stripe uses.

  4. It sends the HTTP POST with a strict 5-second timeout and records the status code and latency.

  5. If the endpoint returns a 5xx or times out, Temporal schedules a retry with exponential backoff and jitter.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Slow endpoints exhausting worker threads

Mitigation

Strict 5-second timeouts and isolated worker pools per customer to avoid head-of-line blocking.

Failure mode 02

A customer outage flooding the retry queues

Mitigation

Disable endpoints that return 5xx errors continuously for more than 24 hours, and email the developer.

Failure mode 03

Server-side request forgery (SSRF)

Mitigation

Reject webhook URLs that resolve to private or internal ranges (127.0.0.0/8, 10.0.0.0/8, 169.254.169.254), checked again at delivery time.

Questions

What teams ask about this design

Retries wait progressively longer (for example 5 s, 1 min, 15 min, 1 h, 6 h, 24 h), with random jitter so thousands of retries don't hit an endpoint at the same instant.

They compute an HMAC-SHA256 of the payload with their signing secret and compare it with the signature header sent with the webhook.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.