Kafka On-Disk Storage: Partition Logs, Segments, and Indexes
At the highest level, Kafka stores data per topic-partition as an ordered, append-only log split into segment files on the broker's local disk. Each partition maps to a directory named <topic>-<partition> (e.g., orders-3), and inside it Kafka writes a sequence of segment files. A segment is the unit of retention, compaction, and recovery. The active segment is the one currently being appended to; older segments are immutable. The name of each segment is its base offset, so 00000000000000000000.log starts at offset 0, and the next segment might start at 00000000000000012345.log. Alongside each .log file are two index files: an .offset index (sparse mapping from offset to byte position) and a .time index (sparse mapping from timestamp to offset). Sparse indexes are key to performance: they are small enough to memory-map, yet allow Kafka to seek close to the target record and then scan forward.
The mechanism matters because it explains Kafka's throughput and recovery behavior. Writes are pure sequential appends to the active segment; reads are sequential scans or positioned by the index. Kafka relies heavily on the OS page cache instead of a JVM heap cache, so the broker can restart quickly because the OS retains hot data. Flushing to disk is configurable: log.flush.interval.messages and log.flush.interval.ms control fsync frequency. In practice, most production clusters rely on replication for durability and let the OS flush lazily; fsyncing every message kills throughput. This is a trade-off: with replication factor 3 and acks=all, you get durability without per-message fsync, but a power-loss event can still lose recently written data that was not yet flushed on all replicas.
Two retention mechanisms operate on these segments. Delete-based retention removes whole segments once they exceed retention.ms or retention.bytes. Compaction keeps the latest record per key indefinitely by rewriting segments and discarding older values for the same key; tombstones (null values) mark deletions. A common misconception is that compaction is synchronous and immediate. It is not. Log cleanup runs asynchronously in the background and is throttled by log.cleaner.threads and related configs. Another misconception is that compacted topics are small; if keys are high-cardinality or tombstones accumulate, the log can stay large because tombstones are retained for delete.retention.ms before being eligible for removal.
Version-dependent flags matter here. In older Kafka versions (pre-2.8/3.0), the behavior of log.roll.ms and segment rolling was tied to log.roll.hours and time-based rolling; more recent versions expose more granular segment configuration and improved tiered storage integration in Kafka 3.x with KIP-405. Also, the KRaft metadata mode introduced in 3.x changes how some internal topics and metadata are stored but does not change the user-topic log format. If you are on a managed service, the provider may abstract or alter retention defaults, so always validate against the actual broker config.
Partition log = ordered append-only sequence of records, split into segments named by base offset.
Segment files include .log (data), .offset index, and .time index; indexes are sparse and memory-mapped.
Retention deletes whole segments; compaction rewrites segments to keep the latest value per key.
Durability comes from replication and acks, not per-message fsync, in most production setups.
Compaction is asynchronous and throttled; high-cardinality keys or tombstone retention can make compacted topics larger than expected.
Tiered storage (KIP-405) in Kafka 3.x offloads older segments to object storage; behavior is version- and provider-dependent.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience