Common Causes of Excessive Consumer Rebalances
Excessive rebalances are one of the most common causes of consumer lag and instability in Kafka. A rebalance occurs when the consumer group's membership changes or when the group coordinator decides that partitions need to be reassigned. During a rebalance, all consumers in the group stop processing and wait for the new assignment, which causes lag to grow. The most common causes are: a consumer crashing or being killed, a consumer failing to send heartbeats within session.timeout.ms, a consumer taking too long between poll() calls and exceeding max.poll.interval.ms, a new consumer joining or leaving the group, and a consumer being fenced because another instance with the same group.instance.id joined. Each of these has a different root cause and a different fix. The first step in diagnosing is to check the consumer group's state and the logs for rebalance events, and to look at the consumer's configuration for the timeout and poll settings.
The mechanism for rebalances is the consumer group protocol. The group coordinator (a broker) manages the group and decides when to trigger a rebalance. Consumers send heartbeats to the coordinator; if a heartbeat is missed for session.timeout.ms, the consumer is removed from the group and a rebalance is triggered. Separately, the consumer must call poll() at least every max.poll.interval.ms; if it does not, the coordinator considers it dead and removes it. This is a common cause of rebalances when processing is slow: the consumer is busy processing a batch and does not return to poll() in time. The fix is to reduce max.poll.records so that each poll returns fewer records, or to move processing to a separate thread, or to increase max.poll.interval.ms if the processing time is legitimately long. A third cause is a consumer that crashes repeatedly, for example because of an OutOfMemoryError or an unhandled exception; the group sees the consumer leave and rejoin, triggering rebalances. A fourth cause is a deployment that rolls consumers one by one; each new instance joining triggers a rebalance. Using static membership (group.instance.id) and a rolling restart strategy can reduce the impact. The trade-off is between session.timeout.ms and rebalance sensitivity: a shorter timeout detects failures faster but triggers more rebalances on transient issues; a longer timeout is more tolerant but slower to detect real failures.
A common mistake is to set max.poll.interval.ms very high to avoid rebalances, which means a stuck consumer is not detected and the group may not make progress. Another mistake is to ignore the consumer's logs, which often show the reason for the rebalance (e.g., "Member consumer-1 has left the group" or "Attempt to heart beat failed"). A third mistake is to blame the broker when the cause is the consumer's processing time. The trade-off is between stability and failure detection. A stable group with long timeouts is good for throughput but slow to recover from real failures; a sensitive group with short timeouts recovers quickly but rebalances on transient issues. The right configuration depends on the workload. For most consumers, max.poll.records should be tuned so that processing a batch takes well under max.poll.interval.ms. Version note: cooperative rebalancing (KIP-429, Kafka 2.4+) reduces the pause during rebalances by allowing consumers to keep their partitions while others are reassigned. Static membership (KIP-345) allows a consumer to rejoin with the same identity and avoid a rebalance if it restarts quickly. Both are important for reducing rebalance impact; check your client version and enable them if available.
A consumer crash or kill triggers a rebalance.
Missed heartbeats beyond session.timeout.ms trigger a rebalance.
Exceeding max.poll.interval.ms between poll() calls triggers a rebalance.
A new consumer joining or leaving the group triggers a rebalance.
Fencing due to duplicate group.instance.id triggers a rebalance.
Tune max.poll.records so processing a batch takes well under max.poll.interval.ms.
Use cooperative rebalancing and static membership to reduce rebalance impact.
Check consumer logs for the specific reason for the rebalance.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience