02 / 05

How should retention be selected for a production topic?

Difficulty: 4/10
Governance, Retention, Runbooks

Selecting Retention for a Production Topic

Retention for a production topic should be selected based on three factors: replay requirements, storage cost, and business needs. The most important is replay requirements: how far back does a consumer need to be able to re-read? If a consumer group can be down for a day and needs to catch up, retention must be at least one day plus a margin. If a consumer needs to rebuild a read model from scratch, retention must cover the entire history the read model needs. If the topic is used for event sourcing or audit, retention may need to be indefinite, which is usually implemented with log compaction or tiered storage. The second factor is storage cost: retention determines how much data is stored, and storage is the largest cost in most Kafka clusters. A 7-day retention on a high-volume topic can cost millions of dollars per year; extending to 30 days multiplies the cost. The third factor is business needs: some topics are used for real-time processing and do not need long retention; others are used for compliance and must be retained for years. The trade-off is between replayability and cost. Longer retention gives more flexibility but costs more; shorter retention is cheaper but limits replay and recovery.

The mechanism for retention in Kafka is the retention policy: cleanup.policy=delete with retention.ms and retention.bytes, or cleanup.policy=compact for keyed topics. For delete-based retention, Kafka deletes whole segments once they exceed the retention time or size. This means the actual retention can exceed the configured value by up to one segment's worth of data. For compacted topics, Kafka retains the latest value per key indefinitely, with tombstones retained for delete.retention.ms. The choice between delete and compact depends on the use case: delete for event streams, compact for changelog or state topics. The trade-off is between simplicity and flexibility. Delete-based retention is simple and predictable; compaction is more complex but retains the latest state per key. For a production topic, the retention policy should be documented and consistent with the topic's purpose. Version note: tiered storage (KIP-405, Kafka 3.x) allows you to offload older segments to object storage, which can dramatically reduce the cost of long retention. With tiered storage, you can set a short local retention (e.g., 1 day) and a long remote retention (e.g., 1 year), getting the benefit of both. If you are on Kafka 3.x, consider tiered storage for long-retention topics. If you are on an older version, you must choose between local retention and cost.

A common mistake is to set retention to a very long period by default, which increases cost without a clear requirement. Another mistake is to set retention too short, which causes data loss when a consumer is down for an extended period. A third mistake is to ignore the storage cost of replication: retention applies to each replica, so the storage cost is multiplied by the replication factor. The trade-off is between the cost of storage and the cost of data loss or replay. A good practice is to start with a retention that covers the maximum expected consumer downtime plus a margin, and to extend it only when there is a clear business requirement. For compliance topics, use tiered storage or a separate archive. Version note: retention is configured per topic and can be overridden at the broker level. The broker-level default is log.retention.hours=168 (7 days). For production topics, always set retention explicitly rather than relying on the default. Also, monitor the actual storage usage per topic to ensure that the retention policy is achieving the expected cost.

javascript
  1. 1

    Select retention based on replay requirements, storage cost, and business needs.

  2. 2

    Cover the maximum expected consumer downtime plus a margin.

  3. 3

    For read model rebuilds, retention must cover the required history.

  4. 4

    For audit or event sourcing, use compaction or tiered storage.

  5. 5

    Storage cost = produce rate * retention * replication factor / compression.

  6. 6

    Delete-based retention deletes whole segments; actual retention can exceed the configured value.

  7. 7

    Tiered storage allows short local retention with long remote retention.

  8. 8

    Set retention explicitly; do not rely on the broker default.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.