Leader Epochs: Preventing Stale Leadership and Ensuring Log Consistency
A leader epoch is a monotonically increasing number that is incremented every time a partition's leader changes. It is used to distinguish between different leadership terms and to prevent a stale leader from causing inconsistencies. When a partition's leader changes, the new leader is assigned a new epoch. Any records or requests that reference an older epoch are considered stale and are rejected or handled appropriately. The leader epoch is useful in three scenarios: detecting stale leaders, truncating follower logs correctly, and ensuring that consumers and producers do not act on stale metadata. Without leader epochs, a broker that was a leader but has been replaced might still think it is the leader, or a follower might have a log that diverges from the new leader's log, and there would be no reliable way to detect and reconcile the divergence. Leader epochs solve this by providing a clear, monotonically increasing identifier for each leadership term. They were introduced in Kafka 0.11 as part of KIP-101, which replaced the older mechanism of truncating logs based on the high watermark.
The mechanism of leader epochs has two parts: the epoch on the leader and the epoch on the follower. When a follower fetches from the leader, it includes its last known leader epoch and the offset of the last record it has from that epoch. The leader compares this with its own log. If the follower's log diverges from the leader's log, the leader uses the epoch information to determine the point of divergence and tells the follower where to truncate. This is more precise than the old HW-based truncation, which could truncate too much or too little. The second part is in the metadata: the controller maintains the leader epoch for each partition, and brokers and clients receive it in metadata responses. A client that sends a produce or fetch request to a broker with a stale epoch receives a NotLeaderOrFollowerException and refreshes its metadata. This prevents a stale leader from accepting writes that would be lost. The trade-off is between precision and complexity. Leader epochs add a small amount of metadata and require the follower to track its epoch, but they eliminate a class of data loss and inconsistency bugs. Version note: leader epochs are part of Kafka 0.11+ and are used in all modern Kafka versions. In KRaft mode, the controller assigns epochs and the mechanism is the same. The older HW-based truncation is deprecated but may still be relevant for very old clients.
A common mistake is to assume that a broker that was a leader is immediately aware that it is no longer the leader. In a distributed system, there is always a window where a broker may still think it is the leader. The leader epoch, combined with the controller's metadata, closes this window: the broker's epoch is compared with the current epoch, and if it is stale, the broker rejects the request. Another mistake is to ignore the leader epoch when debugging log divergence; the epoch information is essential for understanding why a follower truncated its log. A third mistake is to confuse the leader epoch with the leader ID. The leader ID identifies which broker is the leader; the leader epoch identifies which leadership term. The same broker can be the leader in multiple epochs, and different brokers can be the leader in different epochs. The trade-off is between simplicity and correctness. Leader epochs add complexity to the replication protocol, but they are essential for correctness in a system where leaders change. Version note: KIP-101 introduced leader epochs and the truncation based on them. KIP-320 extended this to consumers, allowing them to detect log truncation and handle it gracefully. Modern clients use leader epochs for both replication and consumer fetching.
Leader epoch is a monotonically increasing number incremented on each leader change.
It distinguishes leadership terms and prevents stale leaders from causing inconsistency.
Followers use the epoch to detect log divergence and truncate correctly.
Clients use the epoch to detect stale metadata and refresh.
Leader epoch is different from leader ID; the same broker can lead in multiple epochs.
Introduced in Kafka 0.11 (KIP-101) and extended to consumers in KIP-320.
Without epochs, log truncation based on HW could truncate too much or too little.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience