05 / 05

Broker processes are available but metadata operations fail in KRaft. How would you diagnose it?

Difficulty: 9/10
KRaft, ZooKeeper, Controller quorum

Diagnosing Metadata Failures in KRaft When Brokers Are Healthy

When broker processes are available but metadata operations fail in KRaft, the problem is in the control plane, not the data plane. The brokers can still produce and consume data because the data plane is independent, but metadata operations like creating topics, changing configurations, or updating ACLs fail because the controller quorum is unhealthy. The first step is to separate the two planes: check whether producers and consumers are working (data plane) and whether metadata operations are failing (control plane). The second step is to check the controller quorum's health. Use kafka-metadata-quorum.sh --bootstrap-server <broker> describe --status to see the leader, the high watermark, and the follower lag. If there is no leader, or if the high watermark is not advancing, the quorum is not functioning. The third step is to check the connectivity between the brokers and the controllers. Brokers communicate with the controller quorum on the controller listener; if the network between them is broken, brokers cannot register or receive metadata updates. The fourth step is to check the controller logs for errors: election timeouts, failed appends, or disk issues on the controllers. The fifth step is to check the controller's disk: the metadata log is stored on disk, and if the disk is full or slow, the quorum cannot commit writes. The trade-off is between the data plane and the control plane: a control-plane failure does not stop data flow, but it prevents changes, which can be serious if you need to create topics or reassign partitions.

The mechanism of the failure is that the controller quorum requires a majority to commit metadata writes. If the quorum has lost its majority (e.g., 2 out of 3 controllers are down), it cannot elect a leader or commit writes, and all metadata operations fail. If the quorum has a leader but the leader cannot replicate to a majority (e.g., due to network issues or slow disks on the followers), writes are not committed. If the brokers cannot reach the controllers, they cannot register or receive metadata updates, and they may become isolated. The diagnosis should therefore focus on three things: the quorum's ability to elect a leader, the quorum's ability to replicate and commit, and the brokers' ability to communicate with the quorum. The trade-off is between the size of the quorum and its resilience. A larger quorum tolerates more failures but requires more nodes to be available for a majority; a smaller quorum is more likely to lose its majority. Version note: KRaft's behavior in these scenarios is defined by the Raft protocol. The kafka-metadata-quorum.sh tool and the KRaft metrics are the primary diagnostic tools. In Kafka 3.x, the controller quorum can be monitored separately from the brokers, which makes it easier to isolate control-plane issues.

A common mistake is to assume that the problem is with the brokers because they are the ones reporting the error. In KRaft, metadata operations are handled by the controllers, so the root cause is often in the controller quorum. Another mistake is to ignore the network between brokers and controllers; if the controller listener is not reachable, brokers cannot get metadata. A third mistake is to check only one controller; the quorum's health depends on the majority, so you must check all controllers. The trade-off is between quick fixes and root cause analysis. A quick fix might be to restart a controller, but if the root cause is a disk issue or a network partition, the problem will recur. A thorough diagnosis should identify the root cause and fix it. Version note: in KRaft, the controller quorum is separate from the broker data plane, so you can often recover the control plane without affecting data flow. This is a significant advantage over ZooKeeper-based Kafka, where a ZooKeeper failure could affect both planes. However, a prolonged control-plane outage can still cause issues, such as brokers being unable to register or topics being unable to be created.

javascript
  1. 1

    Separate the data plane (produce/consume) from the control plane (metadata operations).

  2. 2

    Check the controller quorum status with kafka-metadata-quorum.sh.

  3. 3

    A quorum without a leader or with a stalled high watermark cannot commit metadata writes.

  4. 4

    Check connectivity between brokers and controllers on the controller listener.

  5. 5

    Check controller logs for election timeouts, failed appends, and disk issues.

  6. 6

    Check controller disk usage and latency; the metadata log is on disk.

  7. 7

    The root cause is often in the controller quorum, not the brokers.

  8. 8

    KRaft separates control and data planes, so data flow can continue during a control-plane outage.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.