04 / 05

What should a Kafka incident runbook contain?

Difficulty: 8/10
Governance, Retention, Runbooks

What a Kafka Incident Runbook Should Contain

A Kafka incident runbook should contain seven sections: symptoms and impact, initial triage, common failure scenarios, diagnostic commands, recovery procedures, escalation paths, and post-incident actions. The symptoms and impact section describes what the incident looks like from the user's perspective: consumer lag increasing, producers failing, metadata operations failing, or end-to-end latency increasing. The initial triage section provides a checklist to quickly identify the scope: is it one topic, one consumer group, one broker, or the whole cluster? The common failure scenarios section covers the most likely causes: consumer lag, under-replicated partitions, disk full, network partition, rebalance storm, poison message, schema mismatch, and coordinator failure. For each scenario, the runbook should provide diagnostic commands and recovery steps. The diagnostic commands section lists the tools and metrics to check: kafka-consumer-groups.sh, kafka-topics.sh, kafka-metadata-quorum.sh, JMX metrics, and OS-level tools like iostat and vmstat. The recovery procedures section provides step-by-step instructions for each scenario, including rollback and failover. The escalation path section lists who to contact and when. The post-incident actions section covers the follow-up: root cause analysis, remediation, and runbook updates. The trade-off is between completeness and usability. A runbook that is too long is hard to use during an incident; a runbook that is too short is not helpful. The best runbooks are concise, scenario-based, and tested.

The mechanism for using a runbook is to follow it during an incident and to update it after. The runbook should be stored in a location that is accessible during an incident (e.g., a wiki, a shared drive, or a paging tool). It should be versioned and reviewed regularly. The diagnostic commands should be copy-pasteable, with placeholders for the specific cluster and topic. The recovery procedures should be tested in a staging environment before they are needed. The trade-off is between detailed instructions and flexibility. A runbook that prescribes every step is easy to follow but may not fit the specific incident; a runbook that provides principles is flexible but requires more expertise. For most teams, a scenario-based runbook with specific commands and decision points is the right balance. Version note: the diagnostic tools and metrics differ between ZooKeeper-based and KRaft-based Kafka. For KRaft, use kafka-metadata-quorum.sh for controller quorum issues. For ZooKeeper, use zkCli.sh for metadata issues. Update the runbook when you migrate. Also, if you use managed Kafka, the runbook should reference the provider's tools and support channels.

A common mistake is to write a runbook that is never tested. An untested runbook is a liability: it may contain outdated commands or incorrect assumptions. Another mistake is to make the runbook too generic, with no specific commands or thresholds. A third mistake is to forget the post-incident actions: the runbook should be updated after every incident to reflect what was learned. The trade-off is between the cost of maintaining the runbook and the cost of a prolonged incident. A good runbook reduces mean time to recovery (MTTR) and reduces the stress on the on-call engineer. It is an investment that pays off during the worst moments. Version note: the runbook should be aligned with the SLOs and the alerting. If an alert fires, the runbook should tell the engineer what to do. The runbook should also include the escalation path for when the engineer cannot resolve the issue alone. For a mission-critical cluster, the runbook should include the failover procedure and the recovery drills.

javascript
  1. 1

    Seven sections: symptoms, triage, common scenarios, diagnostics, recovery, escalation, post-incident.

  2. 2

    Include specific commands and thresholds, not just principles.

  3. 3

    Cover lag, URP, disk, network, rebalances, coordinators, and schema issues.

  4. 4

    Test the runbook in staging before you need it.

  5. 5

    Update the runbook after every incident.

  6. 6

    Align the runbook with alerts and SLOs.

  7. 7

    KRaft and ZooKeeper have different diagnostic tools; update accordingly.

  8. 8

    An untested runbook is a liability.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.