Why a Consumer Can Process a Record Twice Despite a Single Producer Send
A consumer can process a record twice even when the producer sent it once because Kafka's delivery guarantee is at-least-once, not exactly-once, at the consumer level. The duplication happens in the window between processing the record and committing the offset. The sequence is: the consumer polls a batch of records, processes them (e.g., writes to a database, calls an API), and then commits the offset. If the consumer crashes after processing but before committing, the offset is not advanced. When the consumer restarts, it resumes from the last committed offset, which is behind the processed records, and re-reads and re-processes them. The producer sent the record once, and Kafka stored it once, but the consumer processed it twice because the commit did not happen. This is the fundamental at-least-once behavior: the consumer guarantees that every record is processed at least once, but not exactly once. The duplicate can also happen during a rebalance: if a partition is reassigned to another consumer while the first consumer is processing a batch but has not committed, the new consumer starts from the last committed offset and re-processes the batch. Version note: this behavior is inherent to at-least-once delivery and is not a bug. Kafka's exactly-once semantics (transactions) can eliminate duplicates for consume-transform-produce within Kafka, but they do not eliminate duplicates when the consumer writes to an external system. For external systems, idempotency is required.
The mechanism that causes the duplicate is the offset commit protocol. The consumer's position advances as it reads records, but the committed offset only advances when the consumer calls commitSync or commitAsync. If the consumer uses auto-commit (enable.auto.commit=true), the commit happens periodically in the background, which means the commit can happen before or after processing, depending on the interval. With auto-commit, the consumer may commit an offset for a batch that it has not finished processing, leading to data loss if processing fails; or it may process a batch and crash before the next auto-commit, leading to duplicates. With manual commit (enable.auto.commit=false), the consumer controls when the commit happens, but the window between processing and committing still exists. The trade-off is between commit frequency and duplicate window. Committing after every record minimizes the duplicate window but is expensive; committing after a batch is more efficient but increases the duplicate window. The trade-off is between throughput and duplicate risk. For a high-throughput consumer, batching commits is necessary; for a low-throughput consumer, committing per record is feasible. Version note: Kafka's idempotent producer and transactions reduce duplicates on the producer side and for consume-transform-produce within Kafka, but the consumer-side duplicate window remains for external side effects. Always design the consumer to be idempotent.
A common mistake is to assume that a single producer send means a single consumer process. It does not; the consumer's commit protocol determines the processing guarantee. Another mistake is to commit the offset before processing, which gives at-most-once and can lose records. A third mistake is to use auto-commit without understanding its behavior; auto-commit can cause both duplicates and data loss. The trade-off is between the cost of idempotency and the risk of duplicates. For a system where duplicates are unacceptable, the consumer must be idempotent (using event IDs and a deduplication store) or use Kafka transactions if the output is also Kafka. For a system where duplicates are tolerable, at-least-once is sufficient. Version note: the duplicate window is inherent to distributed systems; it cannot be eliminated entirely without a distributed transaction, which is expensive and complex. The practical approach is to accept at-least-once and make the consumer idempotent. This is the standard pattern in event-driven systems.
At-least-once delivery means a consumer can process a record more than once.
The duplicate window is between processing and committing the offset.
A crash before commit causes re-processing on restart.
A rebalance can cause another consumer to re-process the batch.
Auto-commit can cause both duplicates and data loss.
Committing after every record minimizes duplicates but is expensive.
Kafka transactions eliminate duplicates within Kafka, not for external systems.
Make the consumer idempotent to handle duplicates safely.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience