04 / 05

A cluster has many under-replicated partitions. What should you investigate?

Difficulty: 8/10
ISR and leader election

Find the common factor (one broker, one rack, or cluster-wide), then check broker health, disk, network, replication throttles and recent changes

Under-replicated partitions (URP) mean some followers are not in the ISR, so durability and failover safety are reduced. The first step is to find the pattern, because the cause differs. If every URP involves the same broker, that broker is the problem: it may be down, slow or overloaded. If the URPs follow a rack or availability zone, suspect the network or infrastructure there. If they are spread across all brokers, suspect cluster-wide load, a leader-side bottleneck or a recent change such as a deploy, a reassignment or a traffic spike. I also check whether it is just noise: a reassignment in progress, or a recently restarted broker catching up, causes temporary URPs that are expected.

javascript
  1. 1

    Broker down or restarting: check process health, OOM kills and crash loops. Followers on a restarting broker show as URP until they catch up, which can take long after a big outage.

  2. 2

    Disk: high await, a failing drive, full disk or a slow volume on the follower (or on the leader) makes fetches and appends slow. A single slow disk can hold back every partition it hosts.

  3. 3

    Network: saturated NICs, packet loss, or cross-AZ links limiting replication throughput. Compare inbound produce traffic times (RF minus 1) against available inter-broker bandwidth.

  4. 4

    Broker load: low RequestHandlerAvgIdlePercent or NetworkProcessorAvgIdlePercent means the broker is saturated and cannot serve replica fetches promptly. Look for traffic spikes, too many partitions per broker, or leader skew concentrating load.

  5. 5

    JVM: long GC pauses stall fetcher threads and can also delay heartbeats. Correlate pauses with ISR shrink events.

  6. 6

    Replication tuning: num.replica.fetchers too low for the partition count, replica.fetch.max.bytes too small for large batches, or an over-aggressive replication throttle (leader.replication.throttled.rate) left on after a reassignment.

  7. 7

    Recent changes: partition reassignment, a new high-volume producer, a config change, a deploy or a topic created with a very large partition count often line up exactly with the first URP alert.

  8. 8

    Common mistake: restarting brokers one after another to clear URPs. It adds catch-up load, destroys evidence and can turn URPs into UnderMinIsr and write outages.

  9. 9

    Common mistake: raising replica.lag.time.max.ms to silence the alert. It hides slow followers, lets them hold back acks=all writes, and delays detection of real problems.

  10. 10

    Common mistake: treating every URP as an incident. Check whether a reassignment or restart explains it, and alert on duration and UnderMinIsr rather than on any single URP sample.

  11. 11

    Version note: metric names and the KRaft tooling above are current for modern releases, while ZooKeeper-based clusters needed controller and ZooKeeper health checks as well. Under KIP-966 (early access in 4.0) the reported ISR and eligible-replica details may differ, so verify on your version.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.