Diagnosing Lag on a Subset of Partitions
Lag rising on only two partitions out of many is a strong signal of partition-specific skew. There are three main causes: key skew, consumer skew, and partition-specific broker issues. Key skew means the two partitions are receiving disproportionately more records because the keys that map to them are hot. For example, if the topic is keyed by customer ID and two large customers dominate the traffic, the partitions that hold those customers will have higher lag. Consumer skew means the consumers assigned to those partitions are slower than the others, perhaps because they are on a weaker instance, or because they are doing more expensive processing for those keys, or because they are blocked on a downstream dependency. Partition-specific broker issues mean the broker hosting the leaders for those partitions is slower, perhaps because of disk saturation, network issues, or GC pauses. The first step in diagnosis is to determine which of these is happening by comparing produce rates, consumer processing rates, and broker metrics for the affected partitions versus the healthy ones.
The mechanism for diagnosis is to look at the data, not guess. First, check the produce rate per partition: if the two partitions receive far more records than the others, it is key skew. Second, check the consumer processing rate per partition: if the consumers on those partitions process fewer records per second, it is consumer skew. Third, check the broker metrics for the leaders of those partitions: if the broker has high disk latency, GC pauses, or network saturation, it is a broker issue. Fourth, check the consumer group assignment: if the two partitions are assigned to the same consumer instance, and that instance is slow, it is consumer skew. Fifth, check for rebalances: if the consumer group is rebalancing frequently, processing pauses, and the partitions that were mid-rebalance will show lag. The trade-off is between quick fixes and root cause analysis. A quick fix might be to add more consumers or rebalance the group, but if the root cause is key skew, that will not help because the hot key still maps to one partition. The right fix depends on the cause: for key skew, you may need to change the key or shard the hot key; for consumer skew, you may need to rebalance or scale the consumer instances; for broker issues, you may need to move partitions or fix the broker. Version note: Kafka's consumer group protocol has improved with cooperative rebalancing (KIP-429) in 2.4, which reduces the impact of rebalances, but skew remains a fundamental issue. Tools like kafka-consumer-groups.sh and JMX metrics are the standard way to diagnose.
A common mistake is to assume that lag on two partitions means those partitions are broken. Usually, they are working correctly but are overloaded. Another mistake is to add more consumers without checking the partition count; if the topic has 10 partitions and the group has 10 consumers, adding an 11th consumer will not help because there is no partition for it to take. A third mistake is to ignore the possibility of a hot key. If you see lag on two partitions and the produce rate is high on those partitions, the key distribution is the likely cause. The trade-off is between partitioning and ordering. If you shard the hot key to spread load, you break ordering for that key. If you keep the key and accept the skew, you need to provision the hot partition with more resources or accept higher lag. The right choice depends on whether ordering is required for the hot key. Version note: Kafka does not provide automatic hot partition detection; you need to monitor per-partition metrics and build your own detection. Some managed services provide this, but the responsibility for key design remains yours.
Lag on a subset of partitions usually means skew: key skew, consumer skew, or broker issues.
Compare per-partition produce rates to detect key skew.
Compare per-partition consumer processing rates to detect consumer skew.
Check broker metrics for the leaders of the affected partitions.
Check consumer group assignment to see if the hot partitions are on one instance.
Check for rebalances, which pause processing and cause temporary lag.
Adding consumers does not help if the topic has no unassigned partitions.
Sharding a hot key spreads load but breaks per-key ordering.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience