Executing large partition reassignments safely
I treat a large reassignment as a capacity-managed migration. Before moving data, I estimate replica bytes, validate destination capacity and rack placement, check ISR health, and capture a baseline for broker and client performance. I divide the work into bounded batches, begin with a canary batch, and adjust the movement rate against production SLOs. The goal is not merely to finish quickly; it is to finish without unacceptable client latency or loss of replica safety.
Preflight broker and rack health, destination disk headroom, network capacity, replica sizes, ISR health, and client SLO baselines.
Stage by broker and replica volume, not merely by topic count; a few very large partitions can dominate load.
Start with a canary batch and monitor replication throughput, under-replicated partitions, ISR shrinkage, disk latency, network saturation, request latency, and consumer lag.
Use throttling controls supported by the installed version and adjust the rate based on measured impact. Avoid overlapping maintenance tasks without estimating combined load.
Define stop conditions before starting, such as sustained SLO violations, growing under-replication, destination disk pressure, or repeated ISR loss.
Track each batch's plan, owner, status, verification result, and next action. Aggressive movement finishes sooner but raises contention; conservative movement protects clients but takes longer.
Common mistake: watching only reassignment progress. A move can progress while client traffic degrades or replication safety falls.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience