Compare production and consumption rates, locate where lag accumulates, then check processing time, rebalances, commits, transactions and downstream dependencies
Lag is simply the production rate exceeding the consumption rate, or consumption stalling, so I want to know which and where. I start with the shape of the problem: is lag growing on all partitions or on a few, and did it start with a deploy, a traffic spike or a dependency change? Per-partition lag from the consumer-groups tool tells me whether this is a capacity problem (every partition grows) or a skew or stuck-partition problem (one or a few). Then I compare the topic's incoming rate to the group's consumption rate: if input rose, it may be a pure capacity issue; if input is flat and consumption fell, something got slower.
Processing time: if poll-idle-ratio-avg is near zero, the consumer is saturated doing work. Per-record time multiplied by arrival rate gives the concurrency you need (Little's law). If one handler got slower, profile it: a new database query, a cold cache, a serialization change.
Downstream dependencies: healthy consumers can still be starved by a slow database, API or lock. Check latency, connection pool wait times, retries and rate limits on everything the handler calls.
Partition distribution and skew: if lag is concentrated, look for a hot key, uneven assignment, or one consumer on a slow host. If there are more consumers than partitions, the extras are idle and adding more will not help.
Rebalance churn: frequent rebalances stall consumption even when every member looks healthy. Check rebalance rate, max.poll.interval.ms breaches and pod restarts.
Poison or retry loops: one message failing repeatedly, with in-place retries, blocks its partition. Lag grows on that partition only while the process looks alive and busy.
Commit behavior: lag is measured from committed offsets. If commits are failing, delayed or batched too coarsely, lag can look high or growing while the consumer is actually keeping up. Check commit latency and failures before concluding capacity is the problem.
Transactions: with isolation.level=read_committed, a consumer cannot read past the last stable offset, so one long or hung producer transaction freezes reads on that partition while the log-end offset keeps growing.
Resource limits: CPU throttling in containers, GC pauses, network saturation and disk contention (for broker-side fetch latency) all look like 'healthy but slow'. Check them against pod limits and broker metrics.
Trade-off in the fix: scaling out helps only up to the partition count; beyond that you need more partitions (with a key remapping cost), faster processing, or internal concurrency. Increasing partitions to fix lag without checking skew is a common wrong first move.
Common mistake: restarting the consumers first. It resets symptoms, destroys evidence like current rates and stack traces, and can add rebalance cost; collect metrics and thread dumps before you act.
My order of operations is: rates, per-partition shape, group stability, consumer saturation, downstream health, then the special cases (poison message, open transaction, commit failures). Version note: the lag-monitoring tools and metric names above are standard in recent releases, but under the KIP-848 group protocol some group-state and member output formats differ, so verify the output format on your version.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience