Essential Broker Metrics for Production Kafka Monitoring
A production Kafka broker must be monitored across five categories: disk, network, request latency, replication and ISR, and resource usage. Disk is the most common bottleneck: monitor disk usage per log directory, disk I/O utilization, and disk latency. If disk usage exceeds 80%, you are at risk of a broker going offline; if disk latency spikes, produce and fetch requests slow down. Network is the second most common bottleneck: monitor bytes in and out per second, network utilization, and connection count. Request latency is the most direct measure of broker health: monitor the 99th percentile of produce, fetch, and metadata request latencies. Replication and ISR are the most important for durability: monitor under-replicated partitions, ISR shrinks and expands, and replication lag. Resource usage: monitor CPU, memory, JVM heap, and garbage collection pauses. A broker that is healthy on disk and network but has long GC pauses will still have high request latency.
The mechanism for each signal is important. Disk usage is per log directory; if you have multiple log directories on the same physical disk, they share I/O. Under-replicated partitions (URP) are the single most important durability signal: if a partition has fewer in-sync replicas than its replication factor, a broker failure can cause data loss. Monitor URP count and alert immediately if it is non-zero for more than a short period. ISR shrinks and expands indicate replication problems; frequent shrinks can indicate network issues or slow disks on the follower. Request latency is exposed per request type; a spike in produce latency often indicates disk saturation, while a spike in fetch latency can indicate network or disk. JVM GC pauses are a common cause of latency spikes; monitor pause time and frequency, and tune the heap and GC algorithm. The trade-off is between the number of metrics and the cost of monitoring. Monitoring every JMX metric is expensive and noisy; focus on the signals that map to user-visible impact: latency, durability, and availability. Version note: Kafka exposes metrics via JMX; most production deployments use a JMX exporter to Prometheus or a similar system. In KRaft mode (Kafka 3.x), some metrics related to ZooKeeper are gone, and new controller metrics are available. Check the version's metric documentation when setting up monitoring.
A common mistake is to monitor only broker-level metrics and ignore per-topic and per-partition metrics. A broker can look healthy while one topic has a hot partition that is causing consumer lag. Another mistake is to alert on too many metrics, which leads to alert fatigue and missed real issues. A third mistake is to ignore disk usage until it is critical; disk usage grows gradually, and you should alert at 70% and again at 85% to have time to react. The trade-off is between coverage and noise. A good approach is to define a small set of golden signals (latency, traffic, errors, saturation) per broker and per topic, and alert on those, with deeper metrics available for debugging. Version note: Kafka 3.x with KRaft has different metadata management, so metrics like ActiveControllerCount and ControllerEventManager replace some ZooKeeper-related metrics. Also, tiered storage (KIP-405) introduces new metrics for remote log segments; if you use tiered storage, monitor those separately.
Disk: usage per log dir, I/O utilization, latency; alert at 70% and 85%.
Network: bytes in/out, network utilization, connection count.
Request latency: 99th percentile for Produce, FetchConsumer, FetchFollower.
Replication: UnderReplicatedPartitions (alert if > 0), ISR shrinks/expands.
Resources: CPU, heap, GC pause time; long GC pauses cause latency spikes.
Monitor per-topic and per-partition metrics, not just broker-level.
KRaft replaces ZooKeeper metrics; tiered storage adds remote log metrics.
0-2 years experience
2-5 years experience
5-8 years experience