Use RF=3 across three failure domains, min.insync.replicas=2, acks=all with idempotence, unclean election off, and define the multi-region story honestly
I start from the requirement, not the settings: what is the acceptable data loss (RPO), the acceptable downtime (RTO), the latency budget, and whether the system must keep accepting writes during a failure. For financial events, acknowledged data must not be lost, so consistency wins over availability. The baseline: replication.factor=3 with replicas spread across three racks or availability zones (broker.rack), min.insync.replicas=2, producers with acks=all and enable.idempotence=true, and unclean.leader.election.enable=false. This tolerates the loss of any one broker or zone without losing acknowledged data and without stopping writes, and it refuses writes rather than weakening durability when two replicas are gone.
Then I check the details that are often missed. Internal topics need matching settings: __consumer_offsets and the transaction state log (transaction.state.log.replication.factor and transaction.state.log.min.isr), otherwise offsets or transactions are less durable than the data. Producers need a bounded delivery.timeout.ms and error handling that treats timeouts as unknown outcomes, with event IDs for deduplication. If the pipeline is Kafka-to-Kafka, use transactions and read_committed consumers. For stronger tolerance, RF=5 with min.insync.replicas=3 survives two simultaneous failures at higher latency and cost. Kafka acknowledges after replication to memory-backed page cache on followers rather than fsync, so durability against correlated power loss comes from independent failure domains, not flush settings; forcing fsync is possible but costly and rarely the right tool. For regional disaster recovery, cross-cluster replication (MirrorMaker 2 or a managed equivalent) is asynchronous, so the RPO is greater than zero and consumer offsets need translation; a stretched cluster across regions can give an RPO of zero but pays cross-region latency on every acks=all write and needs a third location for quorum. I would state these numbers to the business instead of promising zero loss everywhere.
Trade-off: RF=3 with min ISR=2 gives one-failure tolerance for both durability and write availability. RF=3 with min ISR=3 survives no failure for writes and is almost always too strict. RF=5 with min ISR=3 tolerates two failures at more cost and latency.
Trade-off: acks=all latency is set by the slowest in-sync follower, and cross-AZ replication adds network time. I accept a few extra milliseconds for durability, and verify the latency budget with a load test including a failure scenario.
Trade-off: consistency over availability. When fewer than min.insync.replicas replicas are alive, writes fail. For payments that is the correct behavior, and upstream systems need a buffer or retry design to cope.
Common mistake: setting acks=all but leaving min.insync.replicas=1, which allows acks=all to degrade to leader-only durability when the ISR shrinks.
Common mistake: protecting the data topics but leaving __consumer_offsets and the transaction log with weaker replication.
Common mistake: enabling unclean leader election to improve availability, which can silently lose acknowledged financial events.
Common mistake: promising zero data loss across regions while using asynchronous mirroring. State the RPO and RTO explicitly, and test failover, including offset translation and consumer cutover.
Operational requirements: alert on UnderMinIsrPartitionCount and ISR shrink rate, rehearse broker and zone failure regularly, use rack-aware assignment, and keep reassignment throttles in the runbook. Encryption, ACLs and audit logging are separate controls but are part of the same platform contract.
Version note: strict handling of min ISR and eligible leader replicas (KIP-966) is evolving, and was early access in Kafka 4.0, so confirm what your release guarantees. KRaft is the only metadata mode from 4.0, and some replication-related defaults have changed across releases, so set the critical ones explicitly.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience