03 / 05

Q43. How would you make a consumer safe when an event can be delivered twice?

Difficulty: 6/10
At-least-once, Exactly-once, Idempotency

Making Consumers Safe Against Duplicate Delivery

If an event can be delivered twice, the consumer must be idempotent: processing the same event more than once must produce the same result as processing it once. There are three main techniques, and they are often combined. The first is an event ID plus a deduplication store. Each event carries a unique ID (a UUID, a composite key, or a Kafka offset plus topic plus partition). The consumer checks whether it has already processed that ID before applying the effect, and records the ID after processing. The deduplication store can be a database table with a unique constraint, Redis with a TTL, or a compacted Kafka topic. The second is idempotent operations: instead of 'increment counter by 1', use 'set counter to value' or 'upsert row with version'. The third is a uniqueness constraint in the target system, such as a primary key or unique index, which rejects duplicate inserts.

The mechanism matters because the choice depends on the target system and the cost of duplicates. For a database, the most robust pattern is to write the event ID and the effect in the same transaction. If the transaction commits, both are recorded; if it rolls back, neither is. On retry, the unique constraint on event ID causes the insert to fail, and the consumer can treat that as already-processed. This is the inbox pattern. For external APIs that are not transactional, idempotency keys are the standard: send a key with the request and let the API deduplicate. For analytics or metrics, duplicates may be acceptable if the aggregation is idempotent, such as a sum of unique user IDs. The key insight is that you cannot make an arbitrary side effect idempotent from the outside; you need cooperation from the target system or a deduplication layer.

A common mistake is to use a simple in-memory set of seen IDs. This fails on restart and does not work across consumer instances in a group. Another mistake is to check for duplicates and then write without a transaction, which leaves a race window where two instances can both check and both write. The trade-off is between storage cost and safety: a deduplication store grows over time and needs TTL or compaction. If you use a compacted Kafka topic for deduplication, you get durability and replayability but add latency. If you use Redis, you get speed but need to manage persistence and eviction. Also note that exactly-once semantics in Kafka can reduce duplicates for consume-transform-produce within Kafka, but they do not eliminate the need for idempotency when writing to external systems. Version-dependent: Kafka's idempotent producer and transactions (0.11+) help, but the consumer-side idempotency is still your responsibility.

javascript
  1. 1

    Idempotency means processing the same event twice has the same effect as processing it once.

  2. 2

    Use event IDs plus a deduplication store (database unique constraint, Redis, compacted topic).

  3. 3

    Prefer idempotent operations (upsert, set) over non-idempotent ones (increment, append).

  4. 4

    Write the event ID and the effect in the same transaction to avoid race windows.

  5. 5

    In-memory deduplication fails on restart and does not work across consumer instances.

  6. 6

    External APIs need idempotency keys; you cannot make an arbitrary side effect idempotent from outside.

  7. 7

    Kafka EOS reduces duplicates within Kafka but does not eliminate the need for consumer-side idempotency.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.