05 / 05

Design operational procedures for a mission-critical Kafka cluster where data loss is unacceptable.

Difficulty: 9/10
Cluster sizing, Replication, Upgrades

Operational Procedures for a Mission-Critical Kafka Cluster

For a mission-critical Kafka cluster where data loss is unacceptable, operational procedures must cover five areas: replication and durability, monitoring and alerting, disaster recovery, controlled change management, and recovery drills. The first area is replication and durability. Use a replication factor of 3 (or higher) with min.insync.replicas=2, and require producers to use acks=all. This ensures that a record is written to at least two replicas before it is acknowledged. Set unclean.leader.election.enable=false so that a broker that is not in the ISR cannot become leader, which prevents data loss at the cost of availability. Distribute replicas across failure domains (racks, AZs) with broker.rack. Use rack-aware replica assignment and verify it. The second area is monitoring and alerting. Monitor UnderReplicatedPartitions and alert immediately if it is non-zero. Monitor ISR shrinks and expands, request latency, disk usage, and network utilization. Set up alerts with burn rates for SLOs. The third area is disaster recovery: a standby cluster in a different region, replicated with MirrorMaker 2 or a similar tool, with a documented failover procedure and regular failover drills. The trade-off is between durability and availability. Requiring acks=all and min.insync.replicas=2 means that if two replicas are unavailable, the producer cannot write, which is a availability hit but protects data. For a mission-critical cluster, data protection takes priority.

The mechanism for controlled change management is to treat every change as a potential incident. Use a change advisory process: every configuration change, upgrade, and topic change goes through review and is applied in a staging environment first. Use canary deployments for producers and consumers. For broker changes, use rolling upgrades one broker at a time with full ISR verification between steps. Never change unclean.leader.election.enable or min.insync.replicas without understanding the durability implications. For topic changes, use kafka-reassign-partitions.sh to move partitions and monitor the reassignment. For schema changes, use a schema registry with FULL_TRANSITIVE compatibility and CI checks. The fourth area is recovery drills: regularly practice broker failure, rack failure, and region failover. The drills should be scheduled, documented, and measured against RTO and RPO. The trade-off is between the cost of redundancy and the cost of downtime. A mission-critical cluster costs more to operate because of the extra replicas, the standby cluster, and the operational overhead. But the cost of data loss or a prolonged outage is far higher. Version note: KRaft mode (Kafka 3.x) changes the metadata management and recovery procedures. If you use KRaft, the controller quorum must be sized and monitored separately. Tiered storage adds another layer of recovery considerations; ensure that remote storage is durable and that recovery procedures account for it.

A common mistake is to set unclean.leader.election.enable=true to improve availability, which allows a broker that is not in the ISR to become leader and lose data. Another mistake is to set min.insync.replicas=1, which means a single replica is enough for a write to succeed, so a broker failure can lose data. A third mistake is to skip recovery drills; a DR plan that has never been tested is not a plan. The trade-off is between availability and durability. Every choice that improves durability reduces availability in some failure scenario. For a mission-critical cluster, the bias should be toward durability. A good practice is to define the RPO (recovery point objective) and RTO (recovery time objective) explicitly, and to design the cluster and procedures to meet them. For example, RPO=0 requires acks=all, min.insync.replicas=2, and unclean.leader.election.enable=false. RTO=15 minutes requires a standby cluster and an automated failover procedure. Version note: MirrorMaker 2 supports active-active replication and offset translation; it is the standard tool for cross-cluster replication. Check the version and the replication topology when designing DR.

javascript
  1. 1

    Use replication factor 3, min.insync.replicas=2, acks=all, unclean.leader.election.enable=false.

  2. 2

    Distribute replicas across racks/AZs with broker.rack and verify assignment.

  3. 3

    Monitor UnderReplicatedPartitions, ISR shrinks, latency, disk, and network; alert on burn rates.

  4. 4

    Use a standby cluster with MirrorMaker 2 and a documented failover procedure.

  5. 5

    Apply all changes through a controlled process with staging and canary deployments.

  6. 6

    Run regular recovery drills for broker, rack, and region failures; measure against RPO/RTO.

  7. 7

    Never set unclean.leader.election.enable=true or min.insync.replicas=1 on a mission-critical cluster.

  8. 8

    KRaft and tiered storage change recovery procedures; update your DR plan accordingly.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.