02 / 05

What are the tradeoffs between automatic and manual offset commits?

Difficulty: 3/10
Commit and replay

Auto-commit is simple but ties commit timing to poll(), so it can lose or duplicate work; manual commits let you commit exactly when processing is truly done

With enable.auto.commit=true (the default), the consumer commits the offsets returned by the previous poll() at most once every auto.commit.interval.ms (default 5 seconds), and it does so inside the next poll() call. When processing is synchronous inside the poll loop, this behaves as at-least-once: by the time poll() commits batch N, batch N has been fully processed. The cost is duplicates: if the consumer crashes after processing records but before the next commit, up to an interval's worth of records are reprocessed after restart.

The dangerous case is when processing is not finished by the time the next poll() runs. If you hand records to a thread pool, or an exception skips the rest of a batch, auto-commit still advances the offset on its timer, so a crash then loses records that were fetched but never processed. That is silent at-most-once behavior. Manual commits fix this by tying the commit to your definition of done. commitSync() blocks until the broker responds and retries retriable errors, so it is safe but adds latency per call. commitAsync() does not block and does not retry, because a retry could overwrite a newer commit with an older offset; you handle failures in a callback. The common pattern is commitAsync in the loop for throughput and a final commitSync on shutdown or revocation. Manual commits still give at-least-once, not exactly-once, so handlers must be idempotent.

javascript
  1. 1

    Trade-off: auto-commit is fine for idempotent, fast, synchronous handlers where a few seconds of replay is harmless. Manual commit is the choice when losing a record costs money or when processing is asynchronous.

  2. 2

    Trade-off: committing per record gives the smallest duplicate window but is slow. Per batch is the usual balance. I size the batch so the replay cost after a crash is acceptable.

  3. 3

    Common mistake: using auto-commit while processing in a separate thread pool. The commit timer has no idea which records are finished, so a crash loses unprocessed work.

  4. 4

    Common mistake: commitAsync with retry logic, which can commit an older offset after a newer one. Use sync for the final commit and keep async fire-and-log.

  5. 5

    Common mistake: forgetting to commit in onPartitionsRevoked. The new owner then replays everything since the last successful commit.

  6. 6

    Common mistake: assuming manual commit means exactly-once. It narrows the loss window to zero but leaves duplicates, so idempotent handling is still required.

  7. 7

    Version note: recent clients refactored consumer internals (the new consumer implementation used with group.protocol=consumer in Kafka 4.0) and changed some commit-timing details such as when auto-commit runs relative to poll(). The overall guarantees are the same, but verify specifics in the release notes of your version.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.