02 / 05

What is the difference between read_committed and read_uncommitted?

Difficulty: 4/10
Transactional processing, Exactly-once, Idempotency

read_committed vs read_uncommitted: Visibility of Transactional Records

read_committed and read_uncommitted are the two values of the consumer configuration isolation.level. They control whether a consumer can see records that were written as part of a transaction that has not yet committed, or that was aborted. With read_uncommitted (the default), the consumer sees all records as soon as they are written to the log, including records from open transactions and records that will later be aborted. With read_committed, the consumer only sees records that are part of committed transactions, and it does not see records from open or aborted transactions. This is implemented by having the broker track the aborted transactions and the last stable offset (LSO), which is the offset up to which all transactions are either committed or aborted. A read_committed consumer can only read up to the LSO.

The mechanism matters because it determines what a consumer can observe and how it affects lag. A read_committed consumer will appear to have lag up to the LSO, not the high watermark, because it cannot read past open transactions. If a transaction is long-running, the LSO does not advance, and the consumer's lag grows even though there is no real backlog of committed data. This is a common source of confusion: consumer lag for a read_committed consumer includes records that are written but not yet committed. Another important detail is that read_committed consumers must filter out aborted records, which adds some overhead compared to read_uncommitted. The broker maintains an aborted transaction index per partition, and the consumer uses it to skip records that belong to aborted transactions. This is why read_committed is slightly more expensive and why it is not the default.

A common mistake is to set read_committed on a consumer that reads from a topic where the producer does not use transactions. In that case, every record is treated as committed, and read_committed behaves almost like read_uncommitted, except for the LSO behavior. Another mistake is to assume that read_committed gives you exactly-once semantics by itself. It does not. It only controls visibility; exactly-once requires the producer to use transactions and the consumer to commit offsets as part of the same transaction. The trade-off is between latency and correctness. read_uncommitted gives lower latency and simpler semantics but can expose records that are later aborted, which is dangerous for downstream systems that cannot handle duplicates or rollbacks. read_committed gives correctness for transactional pipelines at the cost of additional filtering and the LSO constraint. For most non-transactional workloads, read_uncommitted is fine; for consume-transform-produce pipelines with transactions, read_committed is required. Version note: isolation.level was introduced in Kafka 0.11 alongside transactions, and its behavior has been stable since, though KIP-447 improved transaction scalability without changing isolation semantics.

javascript
  1. 1

    read_uncommitted (default) sees all records, including open and aborted transactions.

  2. 2

    read_committed sees only committed records and filters out aborted ones.

  3. 3

    read_committed consumers can only read up to the Last Stable Offset (LSO).

  4. 4

    Consumer lag for read_committed includes records in open transactions, which can look like a backlog.

  5. 5

    read_committed is required for exactly-once pipelines but does not by itself guarantee exactly-once.

  6. 6

    read_uncommitted has lower latency and less overhead but can expose aborted records.

  7. 7

    isolation.level was introduced in Kafka 0.11; semantics have been stable since.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.