04 / 05

How can cooperative rebalancing reduce disruption compared with eager rebalancing?

Difficulty: 8/10
Coordination and rebalancing

Cooperative rebalancing moves only the partitions that must change owners, over multiple rounds, so unaffected consumers keep processing instead of stopping the world

With the eager protocol, every rebalance begins with all members revoking all of their partitions, rejoining, and receiving a fresh assignment. Even when only one partition needs to move, the whole group stops consuming during the rebalance, and consumers with in-memory state or caches lose them. Cooperative rebalancing (incremental, KIP-429, introduced in 2.4 with CooperativeStickyAssignor) changes the contract: members keep the partitions they will retain, and only partitions that must move are revoked. It takes two rounds. In round one, the leader computes the target assignment and each member revokes only the partitions it is losing; the revoked partitions are temporarily unowned. In round two, those freed partitions are assigned to their new owners. Consumers whose partitions did not move never stop processing.

The benefit is largest when the group is big, when processing is stateful (Kafka Streams state stores, local caches), or when rebalances are frequent, such as during rolling deploys. The costs are real: the total time to reach a final assignment can be longer because of the extra round, listener callbacks change meaning (onPartitionsRevoked and onPartitionsAssigned receive only the delta, and onPartitionsLost handles involuntary loss), and the migration needs care because eager and cooperative assignors cannot be mixed in a live group. The standard upgrade is a rolling two-bounce: first deploy a configuration listing both CooperativeStickyAssignor and the current eager assignor, so the group negotiates the common eager protocol; then, once all members are upgraded, remove the eager assignor in a second rolling deploy to switch to cooperative.

javascript
  1. 1

    Trade-off: less interruption per rebalance, but more rounds and a more complex callback model. For small, stateless groups with rare rebalances, the gain may not justify migration risk.

  2. 2

    Common mistake: assuming the default already gives cooperative behavior. Since 3.0 the default list is RangeAssignor then CooperativeStickyAssignor, and the group uses the first protocol every member supports, so it stays eager until Range is removed. Check your client version and config.

  3. 3

    Common mistake: writing listener code that assumes onPartitionsAssigned receives the full assignment. In cooperative mode it is incremental, so use c.assignment() when you need the whole set.

  4. 4

    Common mistake: expecting cooperative rebalancing to fix a rebalance storm. It reduces the cost of each rebalance, not how often they happen. Fix the triggers (poll interval, restarts, timeouts) first.

  5. 5

    Limit: if most partitions move anyway (a big scale-in or total restart), the benefit shrinks. Sticky assignment minimizes movement but cannot avoid it when membership changes drastically.

  6. 6

    Alternative: static membership (group.instance.id) avoids rebalances on quick restarts, and the two are complementary. Kafka Streams has its own task assignment and has used cooperative rebalancing by default in modern versions, but verify for your release.

  7. 7

    Version note: the KIP-848 protocol (GA in Kafka 4.0) moves reconciliation to the broker and is incremental by design, so partition.assignment.strategy and eager versus cooperative stop being the relevant knobs for groups using it. Confirm which protocol your group runs.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.