Skip to content

Reference architecture

High-throughput IoT telemetry pipeline (100k messages per second)

Ingest, process and analyze industrial IoT sensor streams with sub-second dashboard updates, buffering at every hop so a dropped link delays data instead of losing it.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Sustain 100,000+ telemetry events per second, with burst capacity to 500,000 per second
  • 02Under 500 ms from device reading to live dashboard
  • 03Cost-effective storage for petabytes of historical sensor time series
  • 04Anomaly alerts raised within 2 seconds of the triggering reading

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

High-throughput IoT telemetry pipeline (100k messages per second)

Illustrative reference architecture

  1. 01

    Device gateway

    TLS-encrypted MQTT broker terminating device connections

    EMQX / AWS IoT Core

  2. 02

    Streaming buffer

    Partitioned, high-throughput buffer for raw metric packets

    Amazon Kinesis Data Streams

  3. 03

    Stream processor

    Real-time anomaly evaluation and metric windowing

    Apache Flink / Amazon Managed Service for Apache Flink

  4. 04

    Time-series analytics engine

    Columnar storage for years of telemetry, with codecs suited to repetitive sensor data

    ClickHouse Cloud / self-hosted

Subsystem 01

Device gateway

TLS-encrypted MQTT broker terminating device connections

Typical stack

EMQX / AWS IoT Core

Subsystem 02

Streaming buffer

Partitioned, high-throughput buffer for raw metric packets

Typical stack

Amazon Kinesis Data Streams

Subsystem 03

Stream processor

Real-time anomaly evaluation and metric windowing

Typical stack

Apache Flink / Amazon Managed Service for Apache Flink

Subsystem 04

Time-series analytics engine

Columnar storage for years of telemetry, with codecs suited to repetitive sensor data

Typical stack

ClickHouse Cloud / self-hosted

Subsystem 05

Real-time telemetry dashboard

Live monitoring UI for high-frequency metric graphs

Typical stack

Next.js + WebSockets + WebGL

Data lifecycle

End-to-end data flow

  1. Edge sensors publish Protobuf metric payloads to EMQX over mutual-TLS (mTLS) MQTT.

  2. EMQX forwards validated packets into partitioned Kinesis data streams.

  3. Apache Flink evaluates sliding 10-second windows for threshold anomalies and routes alerts to PagerDuty or Slack.

  4. A batch writer inserts metrics into ClickHouse AggregatingMergeTree tables every second.

  5. The Next.js dashboard subscribes to WebSocket feeds to render live sensor states.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Kinesis throttling during traffic storms

Mitigation

Use Kinesis on-demand capacity mode, or hash partition keys evenly and split hot shards automatically.

Failure mode 02

ClickHouse part-count degradation

Mitigation

Buffer writes in memory and insert in bulk (10,000+ rows per batch) to avoid 'too many parts' errors.

Failure mode 03

Device network disconnections

Mitigation

Buffer readings in SQLite on the device and forward them when the connection returns.

Questions

What teams ask about this design

ClickHouse's columnar storage and compression codecs shrink repetitive sensor data substantially, and its vectorized engine aggregates billions of rows quickly. TimescaleDB is a good fit if you want to stay inside PostgreSQL at smaller volumes. We benchmark the candidates on a sample of your own data before choosing.

EMQX runs as a cluster on Kubernetes behind eBPF-based layer-4 load balancers that spread long-lived TCP connections across nodes. Cluster size is load-tested against your projected device count before launch.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.