The ISR is the set of replicas that are caught up with the leader; it defines what acks=all waits for and which replicas can safely become leader
ISR stands for in-sync replicas: the leader plus every follower that is sufficiently caught up. A follower stays in the ISR as long as it has been fully caught up to the leader's log end within replica.lag.time.max.ms. If it falls behind longer than that (slow disk, network trouble, GC pause, broker down), the leader removes it from the ISR; when it catches up again it is added back. The ISR is dynamic and is recorded in the cluster metadata, so it survives leader changes.
It matters for two reasons. Durability: with acks=all the leader acknowledges a write only after every replica currently in the ISR has it, and the high watermark advances only then. Safe election: when a leader fails, the new leader is normally chosen from the ISR, because those replicas hold every acknowledged record. Electing an out-of-sync replica (unclean leader election) can make the cluster available again but silently discards acknowledged writes, so it is disabled by default. The subtle point is that the ISR can shrink to just the leader, at which point acks=all degrades to leader-only durability. That is the job of min.insync.replicas: if the ISR size drops below it, the broker rejects acks=all writes with NotEnoughReplicas rather than accept writes with weak durability.
Trade-off: a stricter min.insync.replicas protects acknowledged data but reduces write availability when brokers fail. With RF=3 and min ISR=2, you tolerate one broker down for writes; two down stops writes.
Trade-off: a short replica.lag.time.max.ms ejects slow followers sooner (smaller acks=all stalls) but makes the ISR flap on brief pauses. A long one tolerates pauses but lets a slow follower hold back every acks=all write.
Common mistake: thinking ISR means all replicas. It is the caught-up subset, and it can be just the leader.
Common mistake: thinking a follower is judged by how many messages it is behind. Modern Kafka judges by time since it was last fully caught up, not by message count.
Common mistake: enabling unclean.leader.election.enable to fix an availability incident without understanding that it trades acknowledged data for availability.
Alert on: IsrShrinksPerSec without matching expands, and UnderMinIsrPartitionCount above zero. Frequent shrink and expand cycles indicate unstable followers.
Version note: replica.lag.time.max.ms default changed from 10 seconds to 30 seconds in Kafka 2.5, and the older replica.lag.max.messages setting was removed long ago, so older articles describe outdated behavior. Verify defaults for your version.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience