Skip to content

Dropped MQTT messages and database bottlenecks at high IoT volume

High-volume IoT telemetry pipeline and MQTT scaling

Stop losing telemetry under load: cluster the MQTT brokers, buffer in a partitioned stream, batch writes into ClickHouse and buffer on devices while they're offline.

Symptoms

Signs your platform has this problem

If several of these sound familiar, the plan below is where we would start.

01

Brokers failing under connection surges

Large numbers of devices reconnecting at once and overwhelming a single MQTT broker.

02

Dropped telemetry during bursts

No streaming buffer, so messages are lost whenever the database slows down.

03

Slow telemetry dashboards

A relational database struggling to aggregate billions of sensor rows.

Remediation plan

How we fix it, step by step

Each phase ends with a measurement, so you can see what changed before the next one starts.

01

Clustered MQTT brokers

EMQX clustered on Kubernetes, sized for your device count and message rate.

02

Partitioned stream buffer

Buffering incoming telemetry in partitioned Amazon Kinesis streams (or Kafka).

03

ClickHouse storage

Batching writes into ClickHouse MergeTree tables built for time-series aggregation.

04

Live dashboards

Pushing current sensor states to dashboards over WebSockets.

Technical checklist

Remediation checklist

What we check before a change goes to production:

  • Cluster the MQTT brokers behind a load balancer
  • Buffer telemetry in a partitioned stream with shard scaling
  • Write to ClickHouse in large batches, not row by row
  • Buffer telemetry on devices (for example in SQLite) for offline periods

What we measure

We take a baseline first and report the same measurements after each change, from your own monitoring — evidence, not promised results.

Throughput
Messages per second sustained in a load test
Message loss
Messages sent vs stored, from sequence numbers
End-to-end delay
Device timestamp to dashboard, at p95

Related service

Custom software

Custom software for Indian SMEs and startups, built around your workflow with an agreed scope, review milestones, clear ownership and documented handover.

Explore Custom software

Questions

Questions about this remediation

Columnar compression and vectorized execution let ClickHouse store billions of sensor rows compactly and aggregate them quickly. A dedicated time-series database can be simpler to run for small fleets.

Devices buffer telemetry locally (for example in SQLite) and upload the backlog when the connection returns, with sequence numbers so nothing is counted twice.

Want an engineer to look at this with you?

Send us the symptoms and any metrics you have. We'll reply within one business day, set up a call and agree what to measure before anything changes.