02 / 05

Q37. What is the difference between retention and log compaction?

Difficulty: 3/10
Logs, Segments, Compaction

Retention vs Log Compaction: Two Different Cleanup Policies

Retention and log compaction are both log cleanup policies, but they solve different problems and operate on different units. Retention (cleanup.policy=delete) deletes whole segments once they age out by time (retention.ms) or exceed a size threshold (retention.bytes). It is coarse-grained: Kafka never deletes individual records, only complete segment files. This means a topic with retention.ms=1 hour will hold data until the oldest segment is entirely older than an hour, so actual retention can exceed the configured value by up to one segment's worth of data. Retention is what you use for event streams where every record matters for a bounded window, such as clickstream or metrics data.

Log compaction (cleanup.policy=compact) is fundamentally different. It treats each record as a key-value pair and retains only the latest value for each key, indefinitely. Older values for the same key are discarded during background compaction. A tombstone (a record with a null value) is written when a key is deleted, and the tombstone itself is retained for delete.retention.ms before being removed. Compaction is what you use for changelog-style data: user profiles, configuration state, CDC streams, or Kafka Streams state stores. The key insight is that compacted topics are not bounded by time; a key that has not been updated in two years still has its latest value available.

The most common mistake less experienced engineers make is assuming compaction is synchronous or immediate. It is not. Compaction runs asynchronously in the background, is throttled by log.cleaner.threads and log.cleaner.backoff.ms, and only kicks in when the dirty ratio (the proportion of un-compacted data in a segment) exceeds min.cleanable.dirty.ratio. This means a compacted topic can temporarily hold multiple values per key. Another common mistake is using compaction for high-cardinality keys, which defeats the purpose because almost every record has a unique key and nothing gets compacted. The trade-off between the two is: retention gives you bounded disk usage and a complete history within a window; compaction gives you unbounded disk usage but a reliable latest-state view with no time bound.

javascript
  1. 1

    Retention deletes whole segments by time or size; compaction rewrites segments to keep the latest value per key.

  2. 2

    Retention is bounded in time; compaction is unbounded in time but bounded by key cardinality.

  3. 3

    Compaction is asynchronous and throttled; it does not happen immediately on write.

  4. 4

    Tombstones mark deletes in compacted topics and are retained for delete.retention.ms before removal.

  5. 5

    Common mistake: using compaction on high-cardinality keys, which results in almost no compaction benefit.

  6. 6

    cleanup.policy=compact,delete combines both: latest value per key, with a time-based floor.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.