03 / 05

How do batching, linger and compression affect producer throughput and latency?

Difficulty: 6/10
Reliability and performance

Larger, fuller, compressed batches raise throughput and cut broker load at the cost of added latency and producer CPU and memory

The producer does not send one record per request. It accumulates records per partition into batches in a buffer and a sender thread ships batches to brokers. Three settings shape this. batch.size is the maximum bytes per partition batch (default 16 KB). linger.ms is how long the producer waits for a batch to fill before sending it anyway. compression.type compresses a whole batch before it is sent. Bigger, fuller batches mean fewer requests, fewer round trips, less per-request overhead on the broker, and better compression ratios, because compression works across many similar records. That is why throughput rises. The cost is latency: with a non-zero linger, a record can wait up to linger.ms before leaving, and a bigger batch means more memory per partition.

The key insight is that these settings only matter if batches actually form. At low traffic, a batch never fills, so linger.ms is pure added latency. At high traffic, batches fill before linger expires and linger.ms barely matters. Spreading traffic across many partitions also hurts batching, because each partition has its own batch, which is why many partitions with low volume per partition produces tiny batches. For compression, lz4 and snappy are cheap on CPU, zstd gives better ratios at moderate CPU, and gzip is the slowest. Compressed batches stay compressed on the broker (when the topic compression.type is producer) and are decompressed only by consumers, so compression reduces network, disk and replication cost too.

javascript
  1. 1

    Trade-off: for latency-sensitive paths (user-facing, trading signals) keep linger low and accept smaller batches. For bulk ingestion, raise linger and batch.size. I tune with measured metrics, not guesses.

  2. 2

    Trade-off: compression saves network, disk and replication but costs producer CPU and some latency. On CPU-bound producers or already-compressed payloads (images, encrypted data) it can hurt.

  3. 3

    Common mistake: raising batch.size alone and expecting gains. Without linger.ms or enough traffic the batch is sent partially filled anyway.

  4. 4

    Common mistake: ignoring buffer.memory. If the producer fills the buffer, send() blocks up to max.block.ms and then throws, so very large batches times many partitions can exhaust it.

  5. 5

    Common mistake: comparing end-to-end latency to linger.ms only. Latency also includes request time, replication with acks=all and consumer fetch settings.

  6. 6

    Version note: the linger.ms default was 0 in older clients, and recent releases changed the default (KIP-1030, Kafka 4.0, to a small non-zero value). Check your client version before assuming either behavior.

  7. 7

    Version note: the default partitioner for null keys moved to sticky and then adaptive strategies (2.4 and 3.3+) specifically to improve batch fill, so older advice about round-robin hurting batching may not apply to your version.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.