Replication keeps multiple copies of each partition on different brokers, with one leader serving traffic and followers copying its log
Replication is how Kafka survives broker failures. Each partition is stored on replication.factor brokers. One replica is the leader and handles all produce requests and, by default, all fetches. The others are followers that continuously fetch from the leader using the same fetch protocol consumers use, appending records to their own logs. If the leader's broker dies, one of the followers that is caught up becomes the new leader and clients carry on. Replication is configured per topic and is applied per partition, so a topic with 12 partitions and replication factor 3 has 36 replicas spread across the cluster.
Two details matter beyond the definition. First, consumers only see records up to the high watermark, the offset replicated to all in-sync replicas, so a record is not visible until it is safely copied. Second, replication is not a backup: a bad write or a delete is replicated to every copy within milliseconds. It protects against hardware and broker loss, not against logical errors, which is why you still need retention policy, access controls and, for disaster recovery, cross-cluster replication. Spreading replicas across racks or availability zones with broker.rack is what makes the copies truly independent.
Trade-off: a higher replication factor improves durability and availability but multiplies storage, network and write cost. RF=3 is the common production default; RF=1 means any broker loss makes the partition unavailable or loses data.
Trade-off: replication is asynchronous by default from the producer's point of view. Whether a write waits for followers depends on acks and min.insync.replicas, not on the replication factor alone.
Common mistake: assuming replication factor 3 means three acknowledged copies. Durability depends on acks=all together with min.insync.replicas.
Common mistake: treating replication as backup. Deletes and corrupt data replicate too.
Common mistake: forgetting the internal topics. __consumer_offsets and the transaction state log have their own replication settings (offsets.topic.replication.factor, transaction.state.log.replication.factor) that need to be set to 3 on a real cluster.
Common mistake: placing all replicas in one rack or zone, so a single failure domain takes out every copy. Set broker.rack.
Version note: ZooKeeper was removed in Kafka 4.0, so replica and leader state is managed by the KRaft controller quorum. Consumers fetching from the closest replica (KIP-392) has been available since 2.4.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience