KRaft Controller Quorum: Raft-Based Metadata Consistency
A KRaft controller quorum maintains metadata consistency using the Raft consensus protocol. The quorum consists of an odd number of controllers (typically 3 or 5). One controller is elected leader, and the others are followers. All metadata writes go through the leader. When a broker or an admin client sends a metadata change (e.g., create a topic, change a config, update an ACL), the leader appends the change to its metadata log and replicates it to the followers. Once a majority of the quorum (a quorum of the quorum, so 2 out of 3 or 3 out of 5) has acknowledged the write, the change is committed. The leader then applies the change to its state machine and notifies the followers, which apply it to their state machines. Brokers consume the metadata log to learn about changes. This ensures that all controllers have the same metadata, and that the metadata is durable as long as a majority of the quorum is available. The use of a majority quorum is what allows the system to tolerate failures: with 3 controllers, one can fail; with 5, two can fail.
The mechanism of Raft has three key parts: leader election, log replication, and commit. Leader election happens when the quorum starts or when the leader fails. A follower that has not heard from the leader within an election timeout becomes a candidate, requests votes from other followers, and if it gets a majority, becomes the new leader. The new leader has the most up-to-date log, which ensures that committed entries are not lost. Log replication happens continuously: the leader sends AppendEntries requests to followers, and followers append the entries to their logs and acknowledge. Commit happens when a majority has acknowledged; the leader then marks the entry as committed and applies it to the state machine. The trade-off is between consistency and availability. Raft is a CP system: it prioritizes consistency over availability. If a majority of the quorum is unavailable, the cluster cannot process metadata writes, and metadata operations fail. This is a deliberate choice: it is better to be unavailable than to have inconsistent metadata. Version note: KRaft's Raft implementation is built into Kafka and uses the __cluster_metadata topic for storage. The quorum size and the election timeouts are configurable. In Kafka 3.x, the controller quorum is separate from the broker data plane, so a broker failure does not affect the quorum unless the controllers are co-located.
A common mistake is to confuse the controller quorum with the broker replication. The controller quorum replicates metadata; the brokers replicate data. They are separate concerns. Another mistake is to run an even number of controllers, which does not improve fault tolerance and can cause ties in elections. A third mistake is to place all controllers in the same failure domain, which defeats the purpose of having a quorum. The trade-off is between latency and durability. A larger quorum gives more fault tolerance but higher latency for metadata writes, because more nodes must acknowledge. A smaller quorum is faster but less resilient. For most clusters, 3 controllers is the right balance; for mission-critical clusters, 5. Version note: KRaft's Raft implementation has been improved with each Kafka release. In Kafka 3.3, KRaft became production-ready; subsequent releases improved performance and observability. Check the version for the latest features and metrics. Also, the controller quorum can be monitored with kafka-metadata-quorum.sh, which shows the leader, the followers, and the replication status.
The KRaft controller quorum uses Raft to replicate metadata across controllers.
One leader handles writes; followers replicate and acknowledge.
A write is committed when a majority of the quorum acknowledges.
Leader election happens when the leader fails; the new leader has the most up-to-date log.
Raft is a CP system: consistency over availability; a majority is required for writes.
Use 3 or 5 controllers; even numbers do not improve fault tolerance.
Place controllers in different failure domains to survive a domain failure.
Monitor the quorum with kafka-metadata-quorum.sh and KRaft metrics.
0-2 years experience
2-5 years experience
5-8 years experience