05 / 05

Kafka is healthy but end-to-end event latency increased tenfold. Build an investigation tree.

Difficulty: 9/10
Production incidents, Consumer lag, Rebalancing

Investigation Tree for a Tenfold Increase in End-to-End Latency

When Kafka itself is healthy but end-to-end latency has increased tenfold, the problem is somewhere in the pipeline outside the broker. The investigation tree should cover the producer, the broker, the consumer, and the downstream systems, in that order. Start with the producer: is the producer experiencing queue buildup? Check the producer's buffer.memory, batch.size, and linger.ms. If the producer is producing faster than it can send, records accumulate in the producer's buffer, and the record's timestamp is earlier than the send time. Check the producer's record-queue-time-avg and record-send-rate metrics. If the producer is slow to send, the cause could be network latency to the broker, broker throttling, or a slow serializer. Next, check the broker: even if the broker is 'healthy' by the usual metrics, check the request latency for produce and fetch, the log flush time, and the page cache hit rate. A broker with high disk latency but low CPU may still look healthy on a dashboard. Check the broker's request queues: if the request handler pool is saturated, requests wait. Then check the consumer: is the consumer processing records slowly? Check the consumer's poll interval, processing time per record, and commit frequency. If the consumer is doing synchronous downstream calls, the latency of those calls adds directly to end-to-end latency. Finally, check the downstream systems: a database, an API, or another service that has slowed down will cause the consumer to slow down, which increases end-to-end latency even though Kafka is fine.

The mechanism for latency in each stage is different. At the producer, latency is the time from when the application calls send() to when the record is acknowledged by the broker. This includes batching time (linger.ms), network time, and broker processing time. At the broker, latency is the time to append the record to the log and replicate it to the ISR. If acks=all, the produce latency includes the time for all in-sync replicas to acknowledge. At the consumer, latency is the time from when the record is fetched to when it is processed and committed. If the consumer commits after processing, the commit latency adds to the end-to-end time. At the downstream system, latency is the response time of the external call. The end-to-end latency is the sum of all these stages. A tenfold increase means one or more stages have slowed down significantly. The trade-off is between latency and throughput: a producer with a large linger.ms has higher throughput but higher latency; a consumer with a large max.poll.records has higher throughput but higher latency per record. During an incident, the first step is to measure each stage to find the bottleneck, not to guess. Version note: Kafka 3.x has improved metrics for request latency and queue time. If you use OpenTelemetry or a tracing system, the trace spans will show you exactly where the time is spent. If you do not have tracing, you need to add timestamps at each stage and correlate them.

A common mistake is to focus only on Kafka because the pipeline uses Kafka. Kafka may be healthy while the producer's buffer is full or the consumer's downstream call is slow. Another mistake is to look at averages instead of percentiles. A tenfold increase in the average latency may be caused by a small number of very slow records that are dragging up the average, or by a uniform slowdown. Check the 99th and 99.9th percentiles, not just the average. A third mistake is to ignore the producer's and consumer's internal queues. If the producer's buffer is full, send() blocks, which increases latency for all records. If the consumer's processing queue is full, poll() is delayed, which increases latency. The trade-off is between adding more instrumentation and the cost of that instrumentation. Tracing every record is expensive; sampling is often necessary. But without tracing, diagnosing a multi-stage latency increase is very difficult. A good approach is to have tracing in place before the incident, with sampling that can be increased on demand. Version note: OpenTelemetry has become the standard for distributed tracing, and Kafka instrumentation libraries are available for producers and consumers. If you use a managed Kafka service, check whether it provides tracing or latency metrics per stage.

javascript
  1. 1

    Check the producer first: buffer full, linger.ms, batch.size, acks.

  2. 2

    Check the broker: request latency, log flush time, request queue, disk latency.

  3. 3

    Check the consumer: processing time, poll interval, commit frequency, rebalances.

  4. 4

    Check downstream systems: database, API, other services that the consumer calls.

  5. 5

    Use percentiles (p99, p99.9), not averages.

  6. 6

    Use distributed tracing to identify the stage with the largest increase.

  7. 7

    Do not assume Kafka is the problem just because the pipeline uses Kafka.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.