Distributing Replicas Across Failure Domains
Replicas should be distributed across failure domains to reduce the risk of correlated loss. A failure domain is any component that can fail independently: a host, a rack, a power supply, a network switch, or an availability zone. If all replicas of a partition are in the same failure domain, a single failure in that domain can take down all replicas simultaneously, causing data loss or unavailability. For example, if all three replicas of a partition are on the same rack and the rack's power fails, the partition is offline and the data is unavailable. If the replicas are on different racks, a single rack failure leaves the other replicas available, and the partition stays online. The same logic applies to availability zones in a cloud environment: spreading replicas across AZs protects against an AZ outage. Kafka has built-in support for rack awareness, which allows you to specify the rack (or AZ) of each broker and have Kafka distribute replicas across racks. This is one of the most important configuration choices for a production cluster.
The mechanism for rack awareness is the broker configuration broker.rack. When set, Kafka's replica assignment algorithm ensures that replicas of a partition are spread across as many racks as possible, up to the replication factor. For a replication factor of 3 and 3 racks, each partition has one replica per rack. If a rack fails, the partitions have at least one surviving replica, and Kafka elects a new leader from the surviving ISR. The trade-off is between durability and latency. Spreading replicas across racks or AZs increases the network latency for replication because the replicas are farther apart. In a single data center, cross-rack latency is low (sub-millisecond); across AZs, it is higher (1-2 ms); across regions, it is much higher (tens of ms). For most workloads, cross-AZ replication is acceptable and is the standard for cloud deployments. Cross-region replication is usually done with MirrorMaker or a similar tool, not with Kafka's internal replication, because the latency is too high. The trade-off is between protection and cost: more failure domains give more protection but may require more network bandwidth and higher latency. Version note: rack awareness has been part of Kafka for many versions. In KRaft mode (Kafka 3.x), the metadata management changes but rack awareness for data replication is unchanged. Cloud providers expose AZs as the natural failure domain; set broker.rack to the AZ name.
A common mistake is to set broker.rack but not verify that the replicas are actually spread. You should check the replica assignment with kafka-topics.sh --describe and confirm that the replicas for each partition are on different racks. Another mistake is to use a replication factor that is too low for the number of failure domains; if you have 3 racks but replication factor 2, you cannot survive a rack failure without losing a replica. A third mistake is to ignore the network implications: if cross-AZ traffic is expensive, you may need to accept higher cost for higher durability. The trade-off is between durability and cost. A replication factor of 3 across 3 AZs gives strong durability but triples the storage and increases cross-AZ traffic. A replication factor of 2 across 2 AZs is cheaper but less durable. The right choice depends on the criticality of the data and the budget. For mission-critical data, replication factor 3 across 3 AZs is the standard. Version note: Kafka's default replica assignment does not guarantee rack awareness unless broker.rack is set on all brokers. If some brokers do not have broker.rack, the assignment may not be rack-aware. Verify the configuration on every broker.
A failure domain is a host, rack, power supply, network switch, or availability zone.
If all replicas are in the same failure domain, one failure can lose all replicas.
Set broker.rack on every broker to enable rack-aware replica assignment.
Kafka spreads replicas across racks up to the replication factor.
Replication factor should be at least the number of failure domains you want to survive.
Cross-AZ replication adds latency but is the standard for cloud deployments.
Cross-region replication is usually done with MirrorMaker, not internal replication.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience