KRaft: Kafka's Modern Metadata Architecture
KRaft (Kafka Raft) is the metadata management architecture that replaces ZooKeeper in Apache Kafka. It was introduced in Kafka 2.8 as an early access feature and became production-ready in Kafka 3.3. In KRaft mode, Kafka manages its own metadata using a Raft-based consensus protocol rather than relying on an external ZooKeeper ensemble. The metadata is stored in a Kafka topic called __cluster_metadata, and a set of controller nodes (which can be the same as brokers or separate) form a quorum that replicates and maintains the metadata. The controller quorum elects a leader, and the leader handles all metadata writes. This is a fundamental change: instead of Kafka brokers communicating with ZooKeeper for metadata, they communicate with the controller quorum. The reason KRaft was introduced is that ZooKeeper was an operational burden and a scalability bottleneck. ZooKeeper is a separate system that must be deployed, monitored, upgraded, and secured independently of Kafka. It has its own consistency model, its own failure modes, and its own limits on the number of watches and znodes. As Kafka clusters grew to thousands of brokers and hundreds of thousands of partitions, ZooKeeper became a bottleneck for metadata operations, particularly controller failover and partition reassignment. KRaft removes the external dependency, simplifies operations, and allows Kafka to scale to millions of partitions.
The mechanism of KRaft is based on the Raft consensus protocol. A KRaft cluster has a set of controller nodes that form a quorum. One controller is elected leader, and the others are followers. All metadata writes go through the leader, which appends them to its log and replicates them to the followers. Once a majority of the quorum has acknowledged the write, it is committed. The metadata log is stored in the __cluster_metadata topic, which is an internal Kafka topic. Brokers register with the controller quorum and receive metadata updates by consuming from the metadata log. The controller quorum can be separate from the broker nodes (dedicated controllers) or co-located (combined mode). For production, dedicated controllers are recommended because they isolate control-plane failures from data-plane failures. The trade-off is between operational simplicity and isolation. Combined mode is simpler to deploy but a controller failure affects a broker, and a broker failure can affect the quorum if too many nodes are lost. Dedicated controllers are more resilient but require more nodes and more configuration. Version note: KRaft became production-ready in Kafka 3.3, and ZooKeeper was deprecated in Kafka 3.5 and removed in Kafka 4.0. If you are starting a new cluster, use KRaft. If you are on an older version, plan your migration. The controller quorum size should be 3 or 5 for fault tolerance; an even number does not improve fault tolerance and can cause split votes.
A common mistake is to assume that KRaft is just a drop-in replacement for ZooKeeper. It is a different architecture with different operational procedures, different metrics, and different failure modes. Another mistake is to run KRaft with a single controller in production, which is a single point of failure. A third mistake is to co-locate controllers and brokers without considering the impact of a broker failure on the quorum. The trade-off is between the simplicity of a single system and the resilience of separation. For a mission-critical cluster, dedicated controllers in separate failure domains are the right choice. For a development cluster, combined mode is fine. Version note: the migration from ZooKeeper to KRaft is not a simple rolling upgrade; it requires a specific migration procedure that involves running both systems in parallel during the transition. Kafka 3.5+ provides migration tooling, but it must be planned carefully. Also, KRaft changes the metadata management for topics, ACLs, and configurations; some ZooKeeper-based tooling may not work, so check your tooling compatibility before migrating.
KRaft replaces ZooKeeper with a Raft-based controller quorum for metadata management.
Metadata is stored in the __cluster_metadata topic and replicated by the quorum.
KRaft was introduced to remove the operational burden and scalability limits of ZooKeeper.
Controllers can be dedicated or co-located with brokers; dedicated is recommended for production.
Use 3 or 5 controllers for fault tolerance; even numbers do not help.
KRaft became production-ready in Kafka 3.3; ZooKeeper was removed in Kafka 4.0.
Migrating from ZooKeeper to KRaft is not a simple rolling upgrade; it requires a specific procedure.
0-2 years experience
2-5 years experience
5-8 years experience