01 / 05

What is consumer lag and why is it important?

Difficulty: 3/10
Lag, Metrics, Tracing

Consumer Lag: The Gap Between Production and Consumption

Consumer lag is the difference between the latest offset in a partition (the log end offset, or LEO) and the last offset committed by a consumer group for that partition. In simple terms, it is the number of records that are available in the topic but have not yet been processed by the consumer. Lag is measured per partition and then aggregated per consumer group. It is the single most important operational metric for a Kafka consumer because it tells you whether the consumer is keeping up with the producer. If lag is stable near zero, the consumer is keeping pace. If lag is growing, the consumer is falling behind and will eventually be unable to catch up. If lag is shrinking, the consumer is catching up. Lag is not just a performance metric; it is a freshness metric. A consumer with high lag is processing stale data, which may be unacceptable for real-time use cases like fraud detection or alerting.

The mechanism that produces lag is straightforward: producers append records to the log at some rate, and consumers read and commit at some rate. If the produce rate exceeds the consume rate, lag grows. The causes of a consume rate that is too low are varied: the consumer may be slow because of expensive per-record processing, the consumer may be blocked on a downstream service, the consumer may have too few instances relative to the number of partitions, or the consumer may be experiencing rebalances that pause processing. Lag is also affected by the number of partitions: a consumer group can have at most one consumer per partition, so if the topic has 10 partitions and the group has 10 consumers, each consumer handles one partition. If the group has 20 consumers, 10 are idle. This means the maximum parallelism of a consumer group is the number of partitions, and lag can grow if the partitions cannot keep up with the produce rate. The trade-off is between partition count and overhead: more partitions allow more consumers but increase broker overhead, rebalance time, and metadata size. Version note: Kafka exposes lag via the kafka-consumer-groups.sh tool and via JMX metrics like records-lag-max. The tool is the standard way to inspect lag, but for production monitoring you should export lag to a metrics system like Prometheus or Datadog. In KRaft mode, lag reporting is unchanged, but the tooling may differ slightly.

A common mistake is to monitor only aggregate lag for a consumer group. Aggregate lag can look fine while one partition has a huge backlog, because the aggregate averages across partitions. You must monitor per-partition lag to detect skew. Another mistake is to alert on a fixed lag threshold without considering the topic's throughput. A lag of 10,000 records on a topic with 1,000 records per second is 10 seconds of delay; the same lag on a topic with 1 record per second is almost 3 hours of delay. The right alert is on lag measured in time, not records, or on the rate of change of lag. A third mistake is to assume that lag is always the consumer's fault. It can also be caused by a producer burst, a broker issue, or a partition skew. The trade-off is between sensitivity and noise. Alerting on any lag above zero will generate false positives during normal bursts; alerting only on very high lag will miss slow degradations. A good approach is to alert on lag that is both above a threshold and growing for a sustained period. Version note: Kafka 3.x introduced the ConsumerGroupCommand improvements and better lag reporting, but the fundamental concept is unchanged.

javascript
  1. 1

    Consumer lag = log end offset minus committed offset for a partition.

  2. 2

    Lag measures how far behind a consumer is; it is a freshness metric, not just performance.

  3. 3

    Lag grows when produce rate exceeds consume rate; causes include slow processing, downstream blocking, too few consumers, or rebalances.

  4. 4

    Maximum consumer parallelism equals the number of partitions.

  5. 5

    Monitor per-partition lag, not just aggregate, to detect skew.

  6. 6

    Alert on lag in time (lag / produce rate), not on a fixed record count.

  7. 7

    Lag can be caused by producers, brokers, or consumers; do not assume it is always the consumer.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.