Suspect a downstream-driven feedback loop: a saturated database, retry amplification and excess consumer concurrency, so test each link and break the loop with backpressure
Healthy brokers point away from Kafka and towards the consumers and what they depend on. Growing lag, rising database load and growing retries together suggest a positive feedback loop: the database slows down, handlers take longer or time out, retries add more load to the same database, which slows it further, and lag grows because consumers are stuck waiting. The first move is to stop the amplification while I collect evidence, not to add capacity. The hypotheses I would test are, in order of likelihood: the database is saturated (connections, locks, slow queries, IOPS) for a reason that preceded the lag, such as a bad query plan, a missing index, a long transaction or a batch job; the consumer's concurrency or retry policy is overwhelming a database that cannot absorb it (immediate retries with no backoff, too many threads, a catch-up burst after a pause); connection pool exhaustion is turning slowness into timeouts and more retries; rebalances caused by exceeding max.poll.interval.ms are redelivering records, so the same work is done several times; a poison record or hot row is causing lock contention; or the idempotency check itself (a dedupe lookup per event) became the expensive query.
Each hypothesis has a discriminating test. Database side: active connections versus pool size, lock waits, slow query log, CPU and I/O saturation, and when each began relative to the first lag increase. Consumer side: per-attempt retry counts, handler latency histograms, pool wait time, poll-idle ratio, rebalance rate and duplicate rate. Little's law gives a sanity check: concurrency equals throughput times latency, so if the database latency went from 5 ms to 500 ms, the same thread count now demands 100 times more connections than the database can serve. Then I run controlled experiments: reduce consumer concurrency or pause partitions and see whether database latency recovers within seconds, which proves the consumers are part of the cause; switch retries to exponential backoff with jitter and a cap; add a circuit breaker so a failing dependency pauses consumption instead of being hammered. Pausing partitions while continuing to poll keeps the consumer in the group, avoiding extra rebalances. Only after the loop is broken do I address the root cause (the slow query, the missing index, capacity) and size concurrency to what the database can sustain.
Trade-off: pausing consumption stops the amplification and protects the database but increases lag immediately. That is the correct trade when the alternative is a collapse, and lag is recoverable because Kafka retains the data.
Trade-off: adding consumers or threads looks like the obvious fix for lag, but if the database is the bottleneck it increases load and makes things worse. Size concurrency to downstream capacity, not to lag.
Trade-off: retries in place keep ordering but block the partition and hold the database busy. Delayed retry topics free the partition but break strict ordering for retried keys, so choose per use case.
Common mistake: immediate retries with no backoff or jitter, and with retries layered at several levels (HTTP client, ORM, application, consumer), which multiplies one failure into dozens of calls.
Common mistake: scaling the database or consumer fleet before finding the trigger. Capacity added during a feedback loop is consumed by retries and the problem returns.
Common mistake: ignoring the rebalance loop. A slow handler exceeds max.poll.interval.ms, the group rebalances, records are redelivered and processed again, adding duplicate load and more slowness.
Common mistake: no timeouts on database calls, so threads block indefinitely, the poll loop stalls and the consumer is removed from the group.
Common mistake: restarting everything first. It resets symptoms, triggers a rebalance and a catch-up burst against a database that is already struggling, and destroys evidence.
Prevention: bulkheads and bounded connection pools, circuit breakers, retry budgets, load shedding, idempotent writes so duplicates are cheap, and alerts on retry rate and pool wait time, not only on lag.
Version note: pause and resume behavior is stable across client versions, but rebalance and heartbeat behavior differ under the KIP-848 protocol (GA in Kafka 4.0), where the poll interval is still a client-side limit. Verify the settings and metrics on your version.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience