Factors That Determine Kafka Cluster Size
Kafka cluster sizing is determined by five factors: throughput, storage, replication, network, and failure tolerance. Throughput is the primary driver: how many bytes per second do you need to produce and consume? A cluster that handles 100 MB/s requires fewer brokers than one that handles 10 GB/s. Storage is the second factor: how much data do you need to retain, and for how long? Storage is a function of produce rate, retention period, and replication factor. Network is the third factor: replication, producer traffic, and consumer traffic all traverse the network, and the network must have enough bandwidth to handle the peak. Replication factor is the fourth factor: a replication factor of 3 triples the storage and roughly doubles the network traffic compared to a replication factor of 1. Failure tolerance is the fifth factor: how many brokers can fail simultaneously without losing availability or durability? The cluster must have enough brokers so that the remaining brokers can absorb the load when some fail. Partitions are also a factor: the number of partitions determines the maximum consumer parallelism and affects broker overhead. A useful rule of thumb is to size for peak throughput plus headroom, not for average throughput.
The mechanism for each factor is important for sizing. Throughput per broker depends on disk, network, and CPU. A modern broker with SSD can handle hundreds of MB/s of produce traffic, but the actual number depends on replication factor, acks, and compression. Storage per broker is limited by disk size and by the number of partitions; too many partitions can cause memory pressure and longer recovery times. Network per broker is limited by the NIC; a 10 Gbps NIC can handle about 1.25 GB/s, but you should not saturate it. Failure tolerance requires that the cluster has enough brokers so that when one fails, the remaining brokers can handle the load and the replicas can be re-replicated. A common guideline is to have at least 3 brokers for replication factor 3, and more for higher availability. The trade-off is between cost and resilience. More brokers give more throughput and resilience but cost more. Fewer brokers are cheaper but have less headroom. The right size depends on the workload and the availability requirements. Version note: Kafka 3.x with KRaft supports larger clusters with faster metadata operations, and tiered storage can reduce local storage requirements by offloading older segments to object storage. These features change the sizing calculus; check the version and features available.
A common mistake is to size for average throughput instead of peak. If the peak is 5x the average, the cluster will be overwhelmed during peak. Another mistake is to ignore the network: a cluster with enough disk and CPU but a saturated network will have high latency and lag. A third mistake is to under-provision failure tolerance: a cluster with exactly the number of brokers needed for the workload has no headroom when a broker fails, and the remaining brokers may not be able to keep up. The trade-off is between over-provisioning and under-provisioning. Over-provisioning costs more but gives resilience and room for growth. Under-provisioning saves money but risks outages. A good practice is to size for peak plus 30-50% headroom, and to monitor utilization to know when to add brokers. Version note: partition count is a key sizing input. More partitions allow more consumers but increase broker memory usage, rebalance time, and recovery time. The guideline of 4,000 partitions per broker is conservative; modern Kafka can handle more, but the exact limit depends on the hardware and the workload. Always test with a realistic partition count.
Five factors: throughput, storage, replication, network, and failure tolerance.
Size for peak throughput plus 30-50% headroom, not average.
Storage = produce rate * retention * replication factor * headroom.
Network must handle producer, replication, and consumer traffic.
Failure tolerance requires enough brokers to absorb load when one fails.
Partitions determine max consumer parallelism and affect broker overhead.
Tiered storage and KRaft change the sizing calculus; check version and features.
0-2 years experience
2-5 years experience
5-8 years experience