Skip to content

Growing Kafka consumer lag, rebalance storms and backlogs

Kafka consumer lag and throughput optimization

Find why consumers fall behind — hot partitions, slow downstream writes, poll timeouts — and fix it so they keep up with your normal load.

Symptoms

Signs your platform has this problem

If several of these sound familiar, the plan below is where we would start.

01

Lag growing by millions of messages

Slow downstream database writes leaving consumer groups hours behind real time.

02

Rebalance storms

Processing that exceeds max.poll.interval.ms, triggering rebalance after rebalance.

03

Hot partitions

One partition saturated while other consumers sit idle.

Remediation plan

How we fix it, step by step

Each phase ends with a measurement, so you can see what changed before the next one starts.

01

Partition and lag analysis

Using lag metrics (Burrow or Prometheus exporters) to find hot partitions and slow consumers.

02

Batched downstream writes

Buffering messages and writing them to the database in bulk multi-row inserts.

03

Poll and timeout tuning

Tuning max.poll.interval.ms and max.poll.records so slow batches don't trigger rebalances.

04

Partition key redesign

Choosing keys with enough entropy to spread load evenly, while keeping ordering where it matters.

Technical checklist

Remediation checklist

What we check before a change goes to production:

  • Audit partition key distribution for hot-partition skew
  • Tune max.poll.records and max.poll.interval.ms
  • Convert single-row database inserts into batched multi-row inserts
  • Alert on consumer lag per group and partition

What we measure

We take a baseline first and report the same measurements after each change, from your own monitoring — evidence, not promised results.

Consumer lag
Messages and seconds behind, per consumer group
Throughput
Messages processed per second per consumer
Rebalances
Consumer group rebalances per day

Related service

Cloud & DevOps

Cloud cost optimization, Kubernetes platforms, and CI/CD that make deploys boring — savings and reliability measured in your dashboards, not our deck.

Explore Cloud & DevOps

Questions

Questions about this remediation

When processing a batch takes longer than max.poll.interval.ms, the group coordinator assumes the consumer has died and rebalances the group, which pauses consumption and can repeat if the cause isn't fixed.

We temporarily increase consumer concurrency up to the partition count, raise fetch batch sizes and batch database writes into larger transactions.

Want an engineer to look at this with you?

Send us the symptoms and any metrics you have. We'll reply within one business day, set up a call and agree what to measure before anything changes.