Why Transactional Records Appear as Gaps to Consumers
When a consumer reads a topic that uses transactions, it may see offsets that appear to have gaps: for example, it reads offset 5, then offset 8, with 6 and 7 missing. This is not data loss. Those offsets contain records that were part of an aborted transaction, and a read_committed consumer skips them. Kafka still writes the records to the log, and the offsets are still consumed in the sequence, but the broker marks the transaction as aborted, and the read_committed consumer filters out those records using the aborted transaction index. The gap is a visibility artifact, not a storage artifact. If you look at the raw log with read_uncommitted, you will see the records; they are just not visible to a read_committed consumer.
The mechanism involves transaction markers. When a transaction commits or aborts, the transaction coordinator writes a control record (a commit marker or abort marker) into each partition that was part of the transaction. These control records are not visible to consumers as data; they are used by the broker to maintain the aborted transaction index and the Last Stable Offset (LSO). A read_committed consumer reads up to the LSO and uses the aborted transaction index to skip records that belong to aborted transactions. This is why gaps appear: the offsets are consumed, but the records are filtered out. Another cause of apparent gaps is the LSO itself. A read_committed consumer cannot read past the LSO, so if there is an open transaction, the consumer will stop at the LSO even though there are records beyond it. Those records are not gaps; they are simply not yet visible.
A common mistake is to treat gaps as corruption or loss and start an incident. The first thing to check is the consumer's isolation.level. If it is read_committed, gaps from aborted transactions are expected. If it is read_uncommitted, gaps should not appear from aborted transactions, but they can appear if the consumer is skipping records due to deserialization errors or if it is using a custom offset management scheme. Another mistake is to assume that aborted transactions are rare. In a pipeline with retries, timeouts, or fencing, aborted transactions can be common, especially during rebalances or restarts. The trade-off is between visibility and correctness: read_committed hides aborted records to give a consistent view, but it makes the stream non-contiguous in terms of offsets. If your downstream system relies on contiguous offsets, you need to handle gaps or use read_uncommitted and filter aborted records yourself. Version note: control records and the aborted transaction index were introduced in Kafka 0.11; behavior has been stable, though KIP-447 improved transaction scalability without changing the marker semantics.
Gaps in offsets for a read_committed consumer are usually records from aborted transactions being filtered out.
Control records (commit and abort markers) are written to the log but are not visible as data.
The aborted transaction index and Last Stable Offset (LSO) drive the filtering.
A read_committed consumer cannot read past the LSO, so open transactions delay visibility.
Gaps are a visibility artifact, not data loss; read_uncommitted would show the records.
Aborted transactions can be common during retries, timeouts, rebalances, or fencing.
Control records and the aborted transaction index were introduced in Kafka 0.11.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience