05 / 05

One Kafka Streams task is much slower than others. How would you diagnose it?

Difficulty: 9/10
Stateful stream processing, State stores, Windowing

Diagnosing a Slow Kafka Streams Task: Skew, State, and Assignment

When one Kafka Streams task is much slower than others, the first hypothesis is partition skew. Tasks correspond to input partitions, and if one partition receives disproportionately more records or larger records, its task will be slower. This is common when the key distribution is skewed: for example, a few high-volume keys dominate the traffic. The first diagnostic step is to compare per-partition input rates and record sizes. If one partition is receiving 10x the traffic of others, that is the cause. The fix is to repartition with a better key or to use a composite key that spreads the load. Another cause of skew is a hot key in a join or aggregation, where one key's state grows much larger than others and slows down lookups and updates.

The second hypothesis is state store performance. If the slow task has a large state store, RocksDB may be doing more disk I/O or compaction than other tasks. Check the state store size, the rate of writes, and RocksDB metrics like compaction time and read latency. A common issue is that the slow task's state store has grown because of a hot key, and RocksDB is spending time on compaction. Another issue is that the state store is on a slow disk or a disk that is shared with other I/O-heavy processes. The third hypothesis is task assignment. If one instance has more tasks than others, or if the slow task is on an instance with fewer resources, that can cause the slowdown. Check the number of tasks per instance and the CPU and memory usage of each instance. If the slow task is on an instance that is also handling other heavy tasks, moving tasks or adding instances can help.

The fourth hypothesis is joins and repartitioning. If the topology includes a join that requires repartitioning, the repartition topic may have skew, and the task that processes the skewed partition will be slower. Check the repartition topic's partition sizes and the join's co-partitioning. A common mistake is to assume that adding more instances will fix a slow task; it will not if the bottleneck is a single partition or a hot key, because tasks cannot be split across instances. The trade-off is between parallelism and skew: you can increase partition count to spread load, but you cannot fix a hot key without changing the key. Another mistake is to ignore backpressure: if the slow task is causing consumer lag, the entire application may be throttled, and the slow task may be a symptom rather than the cause. Version note: Kafka Streams metrics expose per-task processing rates, restore times, and state store sizes; use these to attribute the slowdown. In Kafka 3.x, the task assignment and restoration logic has been improved, but skew remains a fundamental issue.

javascript
  1. 1

    Partition skew is the most common cause: one partition gets more or larger records.

  2. 2

    Hot keys cause large state stores and slow RocksDB compaction.

  3. 3

    Check state store size, write rate, and RocksDB compaction metrics.

  4. 4

    Task assignment imbalance: one instance may have more tasks or fewer resources.

  5. 5

    Repartition topics in joins can also be skewed; check their partition sizes.

  6. 6

    Adding instances will not fix a single hot partition; you must change the key.

  7. 7

    Use per-task metrics to attribute the slowdown; don't assume more instances will help.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.