03 / 05

Which consumer fetch behaviors influence throughput and latency?

Difficulty: 5/10
Polling and processing

Fetch size, broker wait time, per-poll batch size and the consumer's own processing capacity together set the throughput and latency trade-off

Think of the fetch path as a pipeline with two separate dials. The first is how the broker answers a fetch: fetch.min.bytes (default 1) is how much data the broker should accumulate before replying, and fetch.max.wait.ms (default 500) is the longest it will wait to reach that amount. With the defaults, the broker replies as soon as any data exists, which is low latency but small, frequent responses. Raising fetch.min.bytes makes responses larger and cheaper per byte, improving throughput and reducing broker and network load, but on a quiet topic a record can wait up to fetch.max.wait.ms before delivery.

The second dial is how much a response may contain: max.partition.fetch.bytes (default 1 MB) caps data per partition and fetch.max.bytes (default about 50 MB) caps the whole response. Both are soft limits, since a single oversized batch is still returned so the consumer cannot get stuck. Separately, max.poll.records (default 500) controls how many already-fetched records one poll() hands to the application, and does not change how much is fetched from the network. The senior insight is that all of these only matter if the consumer is fetch-bound. Most real consumers are processing-bound: if your handler takes 20 ms per record, no fetch tuning will raise throughput beyond 50 records per second per consumer thread. Little's law applies: required concurrency equals arrival rate times processing time.

javascript
  1. 1

    Trade-off: a latency-sensitive consumer (alerts, user-facing) keeps small waits and small batches. A bulk consumer (ETL, analytics) raises fetch.min.bytes and fetch size to cut requests. I tune from metrics, not from rules of thumb.

  2. 2

    Trade-off: larger fetch sizes mean more memory per consumer, roughly partitions assigned times max.partition.fetch.bytes in the worst case, so a consumer with many partitions can need a lot of heap.

  3. 3

    Common mistake: raising max.poll.records and expecting more network throughput. It only changes batch size handed to the app, and can push you toward the max.poll.interval.ms limit.

  4. 4

    Common mistake: setting fetch.min.bytes high on a low-traffic topic and then being surprised by half-second latency spikes, because the broker waits for fetch.max.wait.ms.

  5. 5

    Common mistake: tuning fetch parameters when the real bottleneck is processing time or too few partitions. Check poll-idle-ratio-avg: if it is near zero the consumer is busy processing, not waiting on fetches.

  6. 6

    Cost lever: follower fetching (client.rack with a rack-aware replica selector on the brokers, KIP-392, since 2.4) lets consumers read from a same-zone replica and cut cross-AZ network charges, at the price of slightly higher replication-lag-bound latency.

  7. 7

    Version note: KIP-74 made the fetch size limits soft in 0.10.1; older advice treating them as hard caps is outdated. Also check defaults on your version, since some fetch-related defaults have been revisited in newer releases.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.