Hot Partitions: How Poor Key Choice Creates Imbalance
A hot partition occurs when one partition receives disproportionately more traffic than others. This happens when the key distribution is skewed: a small number of keys account for a large fraction of the records. For example, if you key by customer and a few customers generate millions of events while most generate hundreds, the partitions that hold those large customers will be hot. The default partitioner hashes the key to a partition, so the skew in keys translates directly to skew in partitions. The consequences are severe: the broker hosting the hot partition has higher CPU, network, and disk usage; the consumer processing that partition becomes the bottleneck for the entire consumer group; and the topic's throughput is limited by the slowest partition. In extreme cases, the hot partition can cause broker overload, replication lag, and even outages.
The mechanism is straightforward: if key K appears 10 million times and all other keys appear 100 times, the partition that K maps to will receive orders of magnitude more records than other partitions. The broker's per-partition resources are not isolated by default, so the hot partition competes with other partitions on the same broker for CPU, disk, and network. On the consumer side, one consumer in the group will be assigned the hot partition and will lag behind while others idle. Adding more consumers does not help because the hot partition cannot be split across consumers. This is the fundamental tension between ordering and parallelism: ordering requires all records for a key to go to the same partition, but that same requirement creates hot spots when keys are skewed. The only ways to mitigate are to choose a key with a more uniform distribution, to split the hot key into sub-keys, or to accept the skew and provision for it.
A common mistake is to use a key that has low cardinality, such as a country code or a status field. With only a few distinct values, the hash will map them to a few partitions, and those partitions will be hot. Another mistake is to use a key that is naturally skewed, such as a tenant ID in a multi-tenant system where one tenant dominates. A third mistake is to use a null key and assume round-robin distribution; round-robin distributes evenly, but it gives no ordering, so it is not a solution when ordering is required. The trade-off is between ordering and balance. If you need per-customer ordering, you must key by customer, and you must accept that a large customer creates a hot partition. The mitigation is to detect the skew early, monitor per-partition throughput, and design for it. Version note: Kafka does not provide automatic hot-partition detection or mitigation; you need to monitor per-partition metrics and alert on imbalance. Some managed services provide partition-level metrics, but the responsibility for key design remains yours.
A hot partition occurs when key distribution is skewed and one partition gets most of the traffic.
Low-cardinality keys (country, status) map to few partitions and cause hot spots.
Large tenants or customers can dominate a partition even with high-cardinality keys.
Hot partitions limit throughput, cause consumer lag, and can overload brokers.
Adding consumers does not help; the hot partition cannot be split.
Mitigations: better key, sub-key sharding (breaks ordering), or provision for the skew.
Kafka does not auto-detect hot partitions; you must monitor per-partition metrics.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience