Estimating Storage and Network for a Topic
Estimating storage and network for a topic requires four inputs: the produce rate (bytes per second or records per second with average record size), the retention period, the replication factor, and the compression ratio. The raw storage is the produce rate multiplied by the retention period. For example, 100 MB/s for 7 days is 100 * 86400 * 7 = 60.5 TB. Multiply by the replication factor to get the total raw storage across the cluster: with replication factor 3, that is 181.4 TB. If compression is used, divide by the compression ratio; for example, a 3:1 compression ratio reduces 181.4 TB to 60.5 TB. Add 30% headroom for safety and operational overhead, giving about 78.6 TB. For network, the producer traffic is the produce rate: 100 MB/s. The replication traffic is (replication factor - 1) times the produce rate: 2 * 100 = 200 MB/s. The consumer traffic depends on how many consumers there are and whether they read all the data; if one consumer reads everything, it is 100 MB/s; if multiple consumer groups read, it multiplies. The total network per broker depends on how the partitions are distributed. The trade-off is between precision and simplicity. This estimate is a starting point; real-world usage may be higher because of overhead, index files, and uncompressed segments. Version note: Kafka's log segment files include the .log, .index, and .timeindex files; the index files add a small percentage to the total storage. If you use tiered storage, older segments are offloaded to object storage, which changes the local storage requirement.
The mechanism for storage is that Kafka appends records to segment files and retains them according to the retention policy. The retention policy can be time-based (retention.ms) or size-based (retention.bytes). If both are set, the smaller limit applies. The storage estimate must account for the fact that segments are deleted in whole units, so the actual storage may be slightly higher than the estimate. For network, the mechanism is that producers send to the leader, and followers fetch from the leader. The leader's network traffic is the produce rate plus the replication fetch rate (if the leader is also a follower for other partitions) plus the consumer fetch rate. The follower's network traffic is the replication fetch rate plus the consumer fetch rate if it becomes a leader. The network estimate must account for the fact that a broker may be a leader for some partitions and a follower for others, so its total traffic is the sum. The trade-off is between over-provisioning and under-provisioning. Over-provisioning costs more but gives headroom; under-provisioning saves money but risks saturation. A good practice is to estimate, then add 30-50% headroom, and monitor actual usage to adjust. Version note: if you use compression, the produce rate in bytes is the compressed size, but the network traffic is also compressed. The compression ratio depends on the data; JSON compresses well (5-10x), Avro compresses less (2-3x). Measure the actual compression ratio with your data.
A common mistake is to forget the replication factor when estimating storage. The raw storage is multiplied by the replication factor, so a replication factor of 3 triples the storage. Another mistake is to ignore compression; if you compress, the storage is reduced, but the CPU cost increases. A third mistake is to estimate network based only on producer traffic; replication and consumer traffic can be larger. The trade-off is between the accuracy of the estimate and the effort to produce it. A simple estimate using the formulas above is usually within 20-30% of actual usage. A more accurate estimate requires measuring the actual record size, compression ratio, and consumer behavior. For a new topic, start with the simple estimate and refine after monitoring. Version note: Kafka 3.x with tiered storage can offload older segments to object storage, which reduces local storage but adds remote read latency. If you use tiered storage, the local storage estimate should cover only the hot data, and the network estimate should include remote fetch traffic.
Storage = produce rate * retention * replication factor / compression ratio * headroom.
Network = producer traffic + replication traffic + consumer traffic.
Replication traffic = produce rate * (replication factor - 1).
Consumer traffic multiplies by the number of consumer groups.
Add 30-50% headroom for safety and operational overhead.
Measure actual record size and compression ratio for accuracy.
Tiered storage reduces local storage but adds remote fetch traffic.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience