Debugging Unexpected Retention on a Compacted Topic
When a compacted topic retains records longer than expected, the first thing to internalize is that compaction is asynchronous and best-effort. Unlike retention, which deletes whole segments on a schedule, compaction runs in the background and only when certain conditions are met. So the answer is usually not a single config but a combination of compaction lag, key cardinality, tombstone handling, and segment-level dirty ratio. Start by confirming what 'longer than expected' means: is it that old values for the same key are still present, or that tombstones have not been removed, or that the topic is simply larger than the latest-state size? Each points to a different cause.
The first thing to inspect is the compaction status of the topic. Kafka exposes cleaner metrics such as kafka.log:type=LogCleanerManager,name=cleanable-ratio and per-topic uncleanable-partitions-count. If min.cleanable.dirty.ratio is set high (default 0.5), compaction only kicks in when at least half the segment is dirty. On a low-throughput topic, that threshold may never be reached, so old values linger. Reducing min.cleanable.dirty.ratio or increasing segment.ms can help. The second thing is key cardinality. If most records have unique keys, there is nothing to compact and the topic grows like a normal log. This is a design issue, not a config issue. The third is tombstones. A tombstone (null value) is retained for delete.retention.ms (default 24 hours) before it is eligible for removal. If consumers need time to see the delete, this is intentional; if not, lower it. But be careful: lowering it too much can cause consumers that are behind to miss the deletion and resurrect the key.
Also inspect segment-level behavior. Compaction operates on closed segments, not the active segment. If segment.ms or segment.bytes is large, the active segment may hold many old values for a long time before it is closed and becomes eligible for compaction. The active segment is never compacted. So a topic with segment.ms=7 days can have stale values in the active segment for up to 7 days. Finally, check for broker-level issues: log.cleaner.threads may be too low, log.cleaner.backoff.ms may be too high, or the cleaner may be starved of I/O. In Kafka 3.x with tiered storage, compaction behavior on remote segments can differ, and some managed services disable or throttle compaction entirely. Always verify against the actual broker config, not the topic config alone.
Compaction is asynchronous and only runs when min.cleanable.dirty.ratio is exceeded.
High-cardinality keys mean nothing to compact; this is a design issue, not a config issue.
Tombstones are retained for delete.retention.ms before removal; lowering it can cause consumers to miss deletes.
The active segment is never compacted; large segment.ms or segment.bytes delays compaction.
Broker-level cleaner threads and backoff can starve compaction even if topic configs look correct.
Tiered storage and managed services may alter compaction behavior; verify actual broker config.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience