03 / 05

How would you perform a rolling Kafka upgrade with minimal disruption?

Difficulty: 6/10
Cluster sizing, Replication, Upgrades

Rolling Kafka Upgrade with Minimal Disruption

A rolling Kafka upgrade means upgrading brokers one at a time, or in small batches, so that the cluster remains available throughout. The key requirements are: the new version must be compatible with the old version for the duration of the upgrade, the cluster must have enough headroom to handle the load when a broker is offline, and the upgrade must be staged and validated at each step. The first step is compatibility planning. Check the Kafka upgrade documentation for the version you are moving to. Kafka supports rolling upgrades between certain versions; for example, you can upgrade from 2.8 to 3.x with a rolling upgrade, but you may need to do it in stages (2.8 -> 3.0 -> 3.3). If you are moving from ZooKeeper to KRaft, that is a separate migration with its own procedure and is not a simple rolling upgrade. Check the protocol version compatibility: the inter-broker protocol version must be set to the old version during the upgrade and then bumped after all brokers are upgraded. The log message format version should also be handled carefully. The second step is to ensure the cluster has enough headroom. If a broker is offline during the upgrade, the remaining brokers must handle its partitions. Check that no broker is above 60-70% utilization before starting. The third step is to upgrade one broker at a time, or one rack at a time, and wait for the cluster to stabilize (all partitions in ISR, no under-replicated partitions) before moving to the next.

The mechanism for a rolling upgrade is to stop the broker, upgrade the software, and restart it. When the broker stops, its partitions' leaders are moved to other brokers, and the replicas on the stopped broker go out of sync. When the broker restarts, it catches up from the leaders and rejoins the ISR. During this time, the cluster is available but with reduced redundancy. The inter-broker protocol version controls the format of messages between brokers; during the upgrade, it must be set to the old version so that new brokers can communicate with old brokers. After all brokers are upgraded, you can bump the protocol version to the new version. The log message format version controls the format of messages in the log; it can be set per topic and should be bumped after the upgrade to take advantage of new features. The trade-off is between speed and safety. Upgrading one broker at a time is slow but safe; upgrading multiple brokers at once is faster but reduces redundancy and increases the risk of an outage if something goes wrong. For a mission-critical cluster, upgrade one broker at a time and wait for full ISR before proceeding. For a less critical cluster, you can upgrade a rack at a time. Version note: the exact upgrade procedure depends on the source and target versions. For example, upgrading from ZooKeeper-based Kafka to KRaft-based Kafka (Kafka 3.3+) requires a migration procedure that is more involved than a rolling upgrade. Always check the official upgrade documentation for your specific versions. Also, client compatibility: newer brokers support older clients, but very old clients may not be supported by very new brokers. Check the client compatibility matrix before upgrading.

A common mistake is to skip the compatibility check and upgrade directly to the latest version, which can cause brokers to fail to communicate. Another mistake is to upgrade all brokers at once, which causes a full outage if something goes wrong. A third mistake is to bump the inter-broker protocol version before all brokers are upgraded, which causes old brokers to reject new messages. The trade-off is between upgrade speed and risk. A slow, staged upgrade is safer but takes longer; a fast upgrade is riskier but gets you to the new version sooner. For production clusters, always favor safety. A good practice is to test the upgrade in a staging environment that mirrors production, and to have a rollback plan. If the upgrade fails, you should be able to roll back to the previous version by stopping the new broker and restarting the old version. This requires that the log format is compatible and that you have not bumped the protocol version. Version note: Kafka 3.x introduced KRaft as a production-ready alternative to ZooKeeper. If you are on ZooKeeper, you can upgrade within ZooKeeper-based versions with a rolling upgrade, and migrate to KRaft separately if desired. The KRaft migration is not a rolling upgrade; it requires a specific procedure.

javascript
  1. 1

    Check version compatibility and upgrade path before starting.

  2. 2

    Set inter.broker.protocol.version and log.message.format.version to the old version during the upgrade.

  3. 3

    Ensure the cluster has headroom (no URP, utilization < 70%).

  4. 4

    Upgrade one broker at a time (or one rack) and wait for full ISR before proceeding.

  5. 5

    After all brokers are upgraded, bump the protocol and message format versions.

  6. 6

    Have a rollback plan; do not bump protocol version until all brokers are upgraded.

  7. 7

    KRaft migration is not a rolling upgrade; it has a separate procedure.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.