01 / 05

What is a Kafka transaction?

Difficulty: 3/10
Transactional processing, Exactly-once, Idempotency

Kafka Transactions: Atomic Writes Across Partitions and Topics

A Kafka transaction is a mechanism that lets a producer write a set of records to multiple partitions and topics atomically, such that either all the records become visible to consumers or none of them do. It also lets the producer commit consumer offsets as part of the same transaction, which is what makes exactly-once consume-transform-produce pipelines possible. Before transactions, the idempotent producer (introduced in Kafka 0.11 alongside transactions) could only guarantee no duplicates within a single partition for a single producer session. Transactions extend that guarantee across partitions and across the consume-process-produce boundary, using a transaction coordinator and a two-phase commit protocol.

The mechanism works like this. The producer is configured with a transactional.id, which is unique and stable per producer instance. On startup, the producer calls initTransactions(), which registers the transactional.id with the transaction coordinator and obtains a producer ID (PID) and epoch. Each transaction begins with beginTransaction(), then the producer sends records to one or more partitions and optionally sends consumer offsets via sendOffsetsToTransaction(). When the producer calls commitTransaction(), the coordinator writes a commit marker to each involved partition and marks the transaction as complete. If the producer crashes or calls abortTransaction(), the coordinator writes abort markers, and read_committed consumers skip the records. The crucial detail is that records are written to the log immediately but are not visible to read_committed consumers until the commit marker is written.

A common mistake is to think that transactions cover external side effects. They do not. A transaction only covers Kafka reads and writes. If your producer calls a database or an API inside the transaction, that side effect is not rolled back if the transaction aborts. Another mistake is to share a transactional.id across multiple running producer instances. The second instance to start will fence the first, causing the first to fail with a ProducerFencedException. This is intentional: it prevents two producers from writing under the same transactional identity and creating duplicates. The trade-off is between latency and atomicity. Transactions add a round trip to the transaction coordinator and require markers to be written, which increases latency compared to a non-transactional producer. For simple single-partition writes where the idempotent producer alone is sufficient, transactions are unnecessary overhead. Version note: transactions were introduced in Kafka 0.11; KIP-447 (Kafka 2.5) improved scalability by allowing a single producer to write to many partitions with better resource management, and Kafka Streams EXACTLY_ONCE_V2 uses that improvement.

javascript
  1. 1

    A Kafka transaction makes writes to multiple partitions and topics atomic: all visible or none visible.

  2. 2

    It can also commit consumer offsets in the same transaction, enabling exactly-once consume-transform-produce.

  3. 3

    The transaction coordinator manages the two-phase commit and writes commit or abort markers.

  4. 4

    Records are written to the log immediately but are hidden from read_committed consumers until commit.

  5. 5

    Transactions do not cover external side effects like database writes or API calls.

  6. 6

    transactional.id must be unique per producer instance; sharing it causes fencing.

  7. 7

    Introduced in Kafka 0.11; KIP-447 in 2.5 improved scalability for many partitions.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.