A rebalance happens whenever group membership, the subscription, or the partition set changes, so partition ownership must be recomputed
A rebalance is the process of redistributing partitions among group members. Conceptually, the group's job is to keep the invariant that every subscribed partition has exactly one owner. Anything that breaks the current assignment forces recomputation. The triggers fall into three families: membership changes, subscription changes, and topic metadata changes.
Membership: a consumer joins (new pod, scale-out, restart), leaves cleanly (close() or shutdown), or is removed because the coordinator stopped hearing heartbeats within session.timeout.ms (crash, network partition, long GC or process freeze).
Poll liveness: a member that does not call poll() within max.poll.interval.ms proactively leaves the group, even though its heartbeat thread is healthy.
Subscription: a member changes its subscribed topics, or a regex subscription starts matching a newly created topic.
Partition metadata: partitions are added to a subscribed topic, or a subscribed topic is deleted or created.
Coordinator events: the group coordinator broker fails over, which can force members to rediscover it and rejoin.
In the classic protocol the sequence is: the coordinator marks the group as rebalancing, every member sends JoinGroup, the coordinator picks a group leader, the leader computes an assignment with the configured assignor, and SyncGroup distributes it. The group moves to a new generation, and commits from an old generation are rejected, which is how Kafka fences zombie consumers. With the default eager protocol all members revoke all partitions first, so processing stops group-wide during this window. That is why rebalances are expensive and why frequent ones are a symptom to investigate, not background noise.
Trade-off: a short session.timeout.ms detects crashes quickly but turns pauses and network blips into rebalances. A long one tolerates pauses but delays failover. Static membership (group.instance.id) lets a restarting instance keep its partitions without a rebalance, at the cost of slower removal of a truly dead member.
Common mistake: believing rebalances only happen when a consumer crashes. Scale-out, deploys, topic changes and slow poll loops all trigger them.
Common mistake: confusing session.timeout.ms (heartbeat thread liveness) with max.poll.interval.ms (application thread progress). Either can remove a member.
Common mistake: a regex subscription like orders-.* in a busy cluster. Every new matching topic causes a rebalance for the whole group.
Version note: cooperative rebalancing (2.4+) reduces the interruption but not the number of rebalances. Kafka 4.0 makes the KIP-848 protocol generally available, where the broker drives assignment and members no longer stop the world, and config names differ (for example, session and heartbeat timeouts become broker-controlled). Verify what your cluster and client versions use.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience