Establish the rebalance trigger from coordinator and client evidence, classify it as poll, heartbeat, churn, network or coordinator, then fix the cause rather than the symptom
A group that rebalances every few seconds is not doing useful work: each rebalance interrupts consumption, so lag grows even though everything looks alive. My plan is evidence first, change second. Step one is stabilizing and observing: confirm the symptom from metrics (rebalance-rate-per-hour, rebalance-total, failed-rebalance-total, the group state flapping between Stable and PreparingRebalance) and note when it began and what changed (deploy, config change, autoscaling event, broker maintenance, a new topic matching a regex subscription). If the incident is severe, I reduce variables, for example pause autoscaling and halt rollouts, before investigating, but I avoid restarting everything because that causes more rebalances and destroys evidence.
Step two is finding the trigger reason. The group coordinator broker logs why each rebalance started (new member joined, member left or was removed with its reason, subscription changed), and each member's client log shows why it left or rejoined. Matching these timestamps tells me which of five buckets I am in. Poll interval breaches: a member exceeds max.poll.interval.ms because processing or a downstream call is slow. Heartbeat and session expiry: the heartbeat thread is starved or delayed by GC pauses, CPU throttling, or a network problem. Membership churn: pods are crash-looping, being OOM-killed, failing liveness probes, or flapped by an autoscaler. Network or broker-side problems: connection resets, packet loss, or broker overload making coordinator requests time out. Coordinator conditions: the coordinator broker is overloaded, restarting, or its __consumer_offsets partition leadership is moving. Step three is applying the fix for the bucket found, verifying through the same metrics, and adding an alert so this is caught early next time.
Poll interval fix: reduce max.poll.records so a batch finishes well inside max.poll.interval.ms, add timeouts to every downstream call in the loop, and only then consider raising max.poll.interval.ms. Raising it first hides stuck consumers.
Heartbeat and session fix: look for GC pauses and CPU throttling (container CPU limits are a frequent culprit). Keep heartbeat.interval.ms at roughly one third of session.timeout.ms, and raise the session timeout only within the broker's group.max.session.timeout.ms limit.
Churn fix: stop crash loops, set realistic memory limits, make liveness probes independent of Kafka processing latency, and add stabilization windows to autoscalers so they do not add and remove pods every minute.
Restart mitigation: static membership (group.instance.id, unique per instance and stable across restarts) avoids a rebalance when an instance restarts within the session timeout, and cooperative rebalancing reduces the cost of those that remain. These reduce impact, they do not remove the root cause.
Network and broker fix: check packet loss, connection resets and DNS between consumers and brokers, broker request handler and network thread saturation, and whether the coordinator broker is overloaded or restarting. A coordinator move forces every group on that coordinator to rejoin.
Slow rebalance fix: if rebalances take long, check what onPartitionsRevoked does. Slow commits or state flushes inside the callback stretch every rebalance, and a member that is slow to rejoin stalls the whole group in the eager protocol.
Subscription fix: if a regex subscription matches topics created frequently, tighten the pattern or move to explicit topic lists so topic creation does not trigger group-wide rebalances.
Common mistake: raising session.timeout.ms or max.poll.interval.ms as the first reaction. That can mask the problem and delay real failure detection.
Common mistake: debugging only one consumer's log. The coordinator log shows the group-level view, and the trigger is often a different member than the one you are staring at.
Common mistake: restarting the whole group to recover. All members rejoin together, and you lose evidence and add a large rebalance on top.
Version note: the default session.timeout.ms is 45 seconds in clients from 3.0 (it was 10 seconds before). Under the KIP-848 protocol (GA in 4.0), the broker drives rebalances, session and heartbeat timeouts are broker-controlled, and log messages and metrics differ, so adapt the plan to your protocol and version.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience