SLOs for a Kafka Platform: Publish Latency, Lag, Durability, and Recovery
SLOs for a Kafka platform should cover four categories: publish latency (how long it takes for a produced record to be available), end-to-end latency (how long from production to consumption), consumer lag (how far behind consumers are), and durability and recovery (how likely data is to survive failures and how fast the cluster recovers). Each SLO should be defined as a target percentage over a window, for example: 99.9% of produce requests complete within 50ms over a 30-day window. The SLO must be measurable with the metrics the platform already exports, and it must be tied to user-visible impact. A useful SLO is one that, when breached, tells you something is wrong that matters. For example, a publish latency SLO of 99.9% under 50ms is useful because it maps to the experience of producers; if it is breached, producers are experiencing slow writes. A lag SLO of 99% of consumer groups under 1 minute of lag is useful because it maps to the freshness of data for consumers. The trade-off is between strictness and achievability. An SLO that is too strict will be breached constantly and ignored; an SLO that is too loose will not catch problems. Start with achievable targets and tighten them as the platform matures.
The mechanism for defining and enforcing SLOs has three parts: measurement, error budgets, and alerting. Measurement requires that the platform exports the relevant metrics per topic, per consumer group, and per broker. Publish latency is measured from the produce request to the ack; end-to-end latency is measured from the producer timestamp to the consumer processing timestamp, which requires the producer to set a timestamp and the consumer to compare it. Lag is measured as the difference between the log end offset and the committed offset, converted to time using the produce rate. Durability is measured by UnderReplicatedPartitions and by the time to recover from a broker failure. Error budgets are the allowed amount of SLO violation; if the error budget is exhausted, the platform team stops feature work and focuses on reliability. Alerting should be based on burn rate, not on individual violations: a fast burn rate over a short window indicates a severe issue; a slow burn rate over a long window indicates a gradual degradation. The trade-off is between the number of SLOs and the complexity of the platform. Too many SLOs make it hard to prioritize; too few leave gaps. A good starting set is five to seven SLOs covering the four categories.
A common mistake is to define SLOs that are not measurable with the existing metrics. If you cannot measure it, you cannot enforce it. Another mistake is to define SLOs at the cluster level only, ignoring per-topic and per-consumer-group differences. A cluster-level lag SLO can hide a single topic with a massive backlog. A third mistake is to set SLOs without an error budget or a process for acting on violations. An SLO without consequences is just a dashboard. The trade-off is between centralization and autonomy. A platform team can define platform-level SLOs (broker availability, publish latency), while individual teams define topic-level SLOs (consumer lag for their topics). This layered approach is more work but more meaningful. Version note: Kafka 3.x with KRaft changes some recovery metrics, and tiered storage introduces new latency dimensions for remote reads. SLOs should be reviewed after major version upgrades or architectural changes. Also note that SLOs are not the same as SLAs; an SLA is a contract with penalties, while an SLO is an internal target. Start with SLOs and only create SLAs when you have the operational maturity to meet them.
Cover four categories: publish latency, end-to-end latency, consumer lag, durability and recovery.
Define each SLO as a target percentage over a window, tied to user-visible impact.
Measure per topic, per consumer group, and per broker; avoid cluster-only SLOs.
Use error budgets and burn-rate alerts, not individual violation alerts.
Start with 5-7 SLOs; tighten as the platform matures.
SLOs are internal targets; SLAs are contracts with penalties.
Review SLOs after major upgrades or architecture changes (KRaft, tiered storage).
0-2 years experience
2-5 years experience
5-8 years experience