Investigating Under-Replicated Partitions During Peak Traffic
Under-replicated partitions (URP) during peak traffic are a durability risk: if the in-sync replica count drops below the minimum, a broker failure can cause data loss. The investigation should correlate broker metrics, disk and network health, and replica-fetch behavior. The first step is to identify which partitions are under-replicated and which brokers host their replicas. Use kafka-topics.sh --describe --under-replicated-partitions to list them, and note the leader and ISR for each. If many partitions on one broker are under-replicated, that broker is likely the problem. If the URP set is scattered, the issue may be cluster-wide, such as network saturation or a slow disk on multiple brokers. The second step is to check the broker's disk metrics: if the follower's disk is saturated, it cannot keep up with the leader's writes, and the ISR shrinks. Check disk utilization, disk latency, and log flush time. The third step is to check network metrics: if the network between the leader and follower is saturated or has high latency, replica fetch falls behind. Check network utilization and replica fetch latency. The fourth step is to check the replica fetcher metrics: kafka.server:type=ReplicaFetcherManager,name=MaxLag shows how far behind the follower is, and kafka.server:type=ReplicaFetcherManager,name=MinFetchRate shows the fetch rate. If the fetch rate is low, the follower is not keeping up.
The mechanism for URP is the replication protocol. The leader accepts writes and the followers fetch from the leader. A follower is in the ISR as long as it is caught up within replica.lag.time.max.ms (default 30 seconds in recent versions). If a follower falls behind by more than that, it is removed from the ISR, and the partition becomes under-replicated. During peak traffic, the produce rate increases, and the followers must fetch and write at a higher rate. If the follower's disk cannot handle the write rate, or if the network cannot handle the fetch rate, the follower falls behind. The fix depends on the cause: if the disk is the bottleneck, add more disks or move partitions; if the network is the bottleneck, upgrade the network or reduce cross-broker traffic; if the broker's CPU is the bottleneck, upgrade the instance. The trade-off is between replication factor and cost. A higher replication factor gives more durability but more cross-broker traffic and more disk usage. During peak traffic, the replication traffic can saturate the network, which is why some clusters use rack-aware replication to keep replicas in the same rack and reduce cross-rack traffic. Version note: replica.lag.time.max.ms default changed from 10 seconds to 30 seconds in Kafka 2.5 to reduce false ISR shrinks. If you are on an older version, you may see more URP events during transient slowness. Also, KRaft mode (Kafka 3.x) does not change the replication protocol but changes metadata management; URP metrics are the same.
A common mistake is to ignore URP during peak traffic because it resolves when traffic subsides. This is dangerous because if a broker fails while URP is high, data loss can occur. Another mistake is to increase replica.lag.time.max.ms to hide the problem; this delays ISR removal but does not fix the underlying slowness, and it increases the window of risk. A third mistake is to assume that URP is always a broker problem; it can also be caused by a slow follower, a network issue, or a hot partition that the follower cannot keep up with. The trade-off is between durability and performance. Stricter ISR thresholds give more durability but more URP events; looser thresholds give fewer URP events but more risk. For critical topics, keep the default or stricter and fix the underlying slowness. Version note: Kafka 3.x has improved replica fetch behavior and better metrics for diagnosing URP. Also, if you use tiered storage, the replication behavior for remote segments differs; monitor URP for local segments separately.
Identify which partitions are under-replicated and which brokers host their replicas.
Check the follower's disk metrics: saturation causes ISR shrink.
Check network metrics: saturated network or high latency causes fetch lag.
Check replica fetcher metrics: MaxLag and MinFetchRate.
If many partitions on one broker are URP, that broker is the problem.
Do not hide URP by increasing replica.lag.time.max.ms; fix the underlying slowness.
URP during peak traffic is a durability risk; a broker failure can cause data loss.
Use rack-aware replication to reduce cross-rack traffic if network is the bottleneck.
0-2 years experience
2-5 years experience
5-8 years experience