04 / 05

Q44. When is Kafka exactly-once processing useful for consume-transform-produce pipelines?

Difficulty: 8/10
At-least-once, Exactly-once, Idempotency

When Exactly-Once Processing Pays Off in Consume-Transform-Produce Pipelines

Kafka exactly-once processing is most useful when the entire pipeline is Kafka-to-Kafka: you consume from one or more input topics, transform the records, and produce to one or more output topics, with no external side effects. This is exactly the shape of a Kafka Streams application or a Flink job that reads from Kafka and writes back to Kafka. In this case, transactions can atomically commit both the output records and the consumer offsets, so a failure results in either all outputs being visible and offsets advanced, or none. This eliminates the duplicate-processing window that at-least-once would otherwise create. It is also useful when downstream consumers use read_committed, so they never see aborted records, which avoids the need for downstream deduplication.

The mechanism that makes this work is the transaction coordinator and the two-phase commit. The producer writes records to the output topic partitions as part of a transaction, and also sends the consumer offsets to the transaction. When the transaction commits, the broker marks the records as committed and advances the offsets. If the producer crashes before commit, the transaction is aborted, and the records are not visible to read_committed consumers. This is why the consumer must use isolation.level=read_committed; with read_uncommitted, it can see records that were later aborted. The trade-off is latency and throughput: transactions add a round trip to the transaction coordinator, and the coordinator can become a bottleneck at very high partition counts. For simple pipelines with a single output partition per input, idempotent producer alone may be enough; transactions are needed when multiple partitions or topics are involved and must be atomic.

When is it not useful? When the pipeline writes to an external system. If the transform step calls a database or an API, Kafka transactions cannot cover that side effect, so you still need idempotency or an outbox pattern. In that case, at-least-once plus idempotent writes is usually simpler and faster than trying to force EOS. Another case where EOS is not worth it: when the downstream can tolerate duplicates or when the cost of deduplication is lower than the cost of transactions. For example, a metrics aggregation that is idempotent by nature does not need EOS. A common mistake is to enable EOS everywhere as a default; this adds complexity and can hurt throughput without a corresponding benefit. Another mistake is to forget that EOS requires a stable transactional.id; if two producer instances share the same transactional.id, the second one will fence the first, causing failures. Version note: KIP-447 (Kafka 2.5) improved transaction scalability, and newer clients have better support for transactional consumers, but the fundamental scope remains Kafka-to-Kafka.

javascript
  1. 1

    EOS is most useful for Kafka-to-Kafka consume-transform-produce pipelines.

  2. 2

    Transactions atomically commit output records and consumer offsets.

  3. 3

    Downstream consumers must use isolation.level=read_committed to avoid aborted records.

  4. 4

    EOS does not cover external side effects; use idempotency plus outbox/inbox for those.

  5. 5

    Trade-off: transactions add latency and coordinator load; idempotent producer may suffice for single-partition cases.

  6. 6

    transactional.id must be unique per producer instance; sharing it causes fencing.

  7. 7

    KIP-447 (Kafka 2.5) improved transaction scalability; EXACTLY_ONCE_V2 in Kafka Streams.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.