04 / 05

How would you perform a large partition reassignment without destabilizing a busy Kafka cluster?

Difficulty: 8/10
Safe operations, Reassignment throttling, Production SLO protection

Executing large partition reassignments safely

I treat a large reassignment as a capacity-managed migration. Before moving data, I estimate replica bytes, validate destination capacity and rack placement, check ISR health, and capture a baseline for broker and client performance. I divide the work into bounded batches, begin with a canary batch, and adjust the movement rate against production SLOs. The goal is not merely to finish quickly; it is to finish without unacceptable client latency or loss of replica safety.

javascript
  1. 1

    Preflight broker and rack health, destination disk headroom, network capacity, replica sizes, ISR health, and client SLO baselines.

  2. 2

    Stage by broker and replica volume, not merely by topic count; a few very large partitions can dominate load.

  3. 3

    Start with a canary batch and monitor replication throughput, under-replicated partitions, ISR shrinkage, disk latency, network saturation, request latency, and consumer lag.

  4. 4

    Use throttling controls supported by the installed version and adjust the rate based on measured impact. Avoid overlapping maintenance tasks without estimating combined load.

  5. 5

    Define stop conditions before starting, such as sustained SLO violations, growing under-replication, destination disk pressure, or repeated ISR loss.

  6. 6

    Track each batch's plan, owner, status, verification result, and next action. Aggressive movement finishes sooner but raises contention; conservative movement protects clients but takes longer.

  7. 7

    Common mistake: watching only reassignment progress. A move can progress while client traffic degrades or replication safety falls.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.