03 / 05

What failure can occur if a consumer commits before processing a record?

Difficulty: 5/10
Commit and replay

A crash or error between the commit and the processing permanently skips the record, which is at-most-once delivery

If the offset is committed first, the group has declared 'this record is done' before it actually is. If the consumer crashes, is killed during a deploy, hits an exception, or loses its partitions in a rebalance after the commit but before the processing completes, the next owner resumes from the committed offset, which is already past that record. The record is never processed again. Nothing in Kafka flags this: lag looks healthy, the offsets look perfect, and the data is simply missing downstream. That is at-most-once semantics: each record is processed zero or one times, never twice.

This often happens unintentionally. Auto-commit combined with handing records to another thread commits offsets for records that are still queued. A catch block that logs an exception and continues, while the loop commits regardless, converts a processing failure into a silent skip. Committing at the top of the loop 'to be safe against duplicates' does the same. At-most-once is sometimes a legitimate choice, for example for metrics or sampling where a lost event is cheaper than a duplicate, but it should be a deliberate decision. For business events the standard fix is to process first, commit after, accept at-least-once and make handlers idempotent. For records that can never be processed, the answer is not to skip them silently but to publish them to a dead-letter topic and then commit, so the skip is recorded and recoverable.

javascript
  1. 1

    Trade-off: at-most-once avoids duplicates and the need for idempotent handlers, but loses data on failure. At-least-once never loses data but requires idempotence. For anything representing money, orders or state I choose at-least-once.

  2. 2

    Trade-off: dead-lettering keeps the partition moving and preserves the failed record, but the DLQ needs monitoring and a replay path, otherwise it becomes a silent graveyard.

  3. 3

    Common mistake: swallowing exceptions in the loop and committing anyway. Every failure path must either retry, dead-letter, or stop the consumer, never skip silently.

  4. 4

    Common mistake: assuming a graceful shutdown protects you. A commit before processing is unsafe even on clean restarts, because the in-memory batch is lost when the process stops.

  5. 5

    Common mistake: committing inside a rebalance callback for records not yet processed. On revocation commit only offsets of finished work.

  6. 6

    Detection: gaps are hard to see from lag alone. Reconcile downstream counts against source counts, or compare event IDs or sequence numbers, to find missing records.

  7. 7

    Version note: producer-side idempotence and transactions do not change this consumer-side behavior, and new group protocols do not either. The processing and commit order is always application logic, so it is unaffected by Kafka version.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.