Diagnosing Broker I/O Bottlenecks: Storage, Retention, Replication, or Compaction
When a broker shows high I/O, the goal is to attribute the load to one of four sources: normal storage writes (producer traffic), retention deletion, replication (follower fetch), or compaction. The approach is to correlate broker metrics, per-topic metrics, and OS-level disk metrics. Start with the OS: iostat -x 1 and vmstat 1 tell you whether the disk is saturated (util% near 100), whether reads or writes dominate, and whether the workload is sequential or random. Then map that to Kafka metrics. If writes dominate and correlate with producer byte rate, it is normal storage. If you see periodic spikes that correlate with log deletion, it is retention. If follower fetch rate is high and the broker is a leader, it is replication. If compaction is running, you will see LogCleaner metrics and recopy activity.
For retention, look at the timing of the I/O spikes. Retention deletes whole segments, which is a metadata operation plus file deletion; it should not cause sustained high I/O. But if many segments expire at once, you can see a burst. Check log.retention.check.interval.ms and whether many topics share the same retention window, causing synchronized deletion. For replication, check kafka.server:type=ReplicaFetcherManager metrics and the follower fetch rate. If a broker is a follower and is fetching a lot, that is replication read I/O. If it is a leader and followers are catching up after a restart, you can see a replication storm. Throttle with replica.fetch.max.bytes and follower replication throttles if needed. For compaction, check kafka.log:type=LogCleanerManager,name=cleanable-ratio and kafka.log:type=LogCleaner,name=cleaner-recopy-percent. Compaction rewrites segments, so it shows as both read and write I/O on the same disk, often with a distinct pattern from producer traffic.
The trade-off here is between throughput and isolation. In a multi-tenant cluster, one tenant's compaction or retention burst can starve another tenant's producer traffic. The common mistake is to look only at broker-level metrics and miss per-topic attribution. Kafka exposes per-topic metrics via JMX, but they can be expensive at scale; many teams use a metrics pipeline that aggregates them. Another mistake is to assume that high I/O always means a problem. Kafka is designed to use disk heavily; the question is whether the I/O is causing latency or lag. If producer latency and consumer lag are fine, high I/O may be expected. If not, you need to attribute and throttle. In Kafka 3.x with tiered storage, remote reads can add a new I/O source that is not on local disk, so check tiered storage metrics separately. Also note that KRaft mode changes some internal topic behavior but not the fundamental I/O sources.
Start with OS-level iostat and vmstat to see whether reads or writes dominate and whether the disk is saturated.
Correlate with producer byte rate for normal storage; follower fetch rate for replication.
Retention shows as periodic bursts when segments expire; check retention check interval and synchronized retention windows.
Compaction shows as both read and write I/O on the same disk, with LogCleaner metrics and recopy activity.
Per-topic attribution is essential in multi-tenant clusters; broker-level metrics alone are not enough.
High I/O is not automatically a problem; check producer latency and consumer lag before throttling.
Tiered storage adds remote I/O sources; monitor tiered storage metrics separately in Kafka 3.x.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience