The consumer is removed from its group after max.poll.interval.ms, triggering a rebalance, duplicate processing and failed commits
A consumer in a group must call poll() regularly. The client runs a background heartbeat thread, so a consumer that is busy but alive keeps sending heartbeats and does not hit session.timeout.ms. What it cannot hide is time between poll() calls. If more than max.poll.interval.ms (default 5 minutes) passes without a poll, the client concludes the application thread is stuck or too slow, leaves the group, and the coordinator triggers a rebalance. The partitions are reassigned to other members, who resume from the last committed offset.
The consequences are the interesting part. First, work the slow consumer already did but had not committed will be done again by the new owner, so you get duplicate processing. Second, when the slow consumer finally finishes and tries to commit, the commit fails (CommitFailedException, or a rebalance-in-progress error depending on version) because it no longer owns those partitions. Third, the consumer rejoins on its next poll(), causing another rebalance. If every poll cycle is too slow, this repeats forever: the group flaps, lag grows, and nothing makes progress, a classic rebalance storm. The fix is to make the time per poll batch small and predictable, not just to raise the timeout.
Trade-off: raising max.poll.interval.ms hides the symptom but delays detection of a truly stuck consumer, because a hung member keeps its partitions longer. Lowering max.poll.records bounds batch time without that cost, so I reduce the batch first and raise the interval second.
Trade-off: offloading processing to worker threads decouples polling from processing, but then you must pause partitions for backpressure and track offsets safely. It is the right answer when per-record work is slow and variable.
Common mistake: thinking session.timeout.ms governs this. That timeout is about heartbeats from the background thread; the poll interval is a separate check on the application thread.
Common mistake: ignoring downstream stalls. A database or HTTP call with no timeout can block the polling thread indefinitely, so every call inside the loop needs a timeout shorter than the poll budget.
Common mistake: retrying the failed commit. After removal, the commit can never succeed for those partitions; the correct response is idempotent processing and rejoining.
Version note: the default session.timeout.ms rose from 10s to 45s in Kafka 3.0. Under the KIP-848 protocol (GA in 4.0) session and heartbeat timeouts are controlled by the broker side, while max.poll.interval.ms remains a client-side setting. Verify the details for your version.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience