RPO and RTO: Translating Business Requirements into Kafka DR Design
RPO (Recovery Point Objective) is the maximum amount of data that can be lost, measured in time. If the RPO is 5 minutes, the business can tolerate losing up to 5 minutes of data during a failure. RTO (Recovery Time Objective) is the maximum amount of time the system can be unavailable, measured in time. If the RTO is 15 minutes, the business can tolerate 15 minutes of downtime. In Kafka DR, RPO is determined by replication lag: if the standby cluster is 5 minutes behind the primary, the RPO is at least 5 minutes. RTO is determined by the failover procedure: how long it takes to detect the failure, redirect producers and consumers, and bring the standby cluster into service. These two objectives drive the entire DR design. A low RPO requires near-real-time replication and may require synchronous replication for zero RPO; a low RTO requires automated failover and pre-provisioned standby capacity. The trade-off is between cost and recovery. Lower RPO and RTO require more infrastructure, more network bandwidth, and more automation, all of which cost more. The business must decide what it is willing to pay for.
The mechanism for meeting RPO in Kafka is replication lag management. MirrorMaker 2 replicates asynchronously, so the lag depends on the network bandwidth, the number of partitions, and the throughput. To reduce lag, you can increase the number of MM2 tasks, increase the network bandwidth, and reduce the amount of data replicated (by selecting only critical topics). For RPO=0, asynchronous replication is not enough; you need synchronous replication, which Kafka does not support natively across regions. Some architectures use a synchronous dual-write pattern, where the producer writes to both clusters and waits for both to acknowledge. This gives zero RPO but adds latency and complexity, and it requires handling partial failures (one write succeeds, the other fails). The mechanism for meeting RTO is failover automation. A manual failover takes minutes to hours; an automated failover takes seconds to minutes. Automated failover requires health checks, a decision mechanism (e.g., a consensus system), and a way to redirect producers and consumers to the standby cluster. DNS-based failover is common: the producer and consumer use a DNS name that points to the primary cluster; on failover, the DNS record is updated to point to the standby. The trade-off is between automation complexity and recovery speed. More automation reduces RTO but adds the risk of false positives (failing over when the primary is actually healthy). Version note: MirrorMaker 2 provides heartbeat topics and checkpoint topics that can be used to monitor replication lag and automate failover. Some managed Kafka services provide built-in multi-region replication and failover. Check the features of your platform.
A common mistake is to define RPO and RTO without measuring them. A DR plan that has never been tested is not a plan. You should run regular failover drills and measure the actual RPO and RTO, then compare them to the targets. Another mistake is to assume that asynchronous replication gives zero RPO; it does not. A third mistake is to set an aggressive RTO without the automation to meet it; a manual failover cannot meet a 5-minute RTO. The trade-off is between the cost of DR and the cost of downtime. For a financial platform, the cost of downtime is usually high enough to justify a low RPO and RTO. For an internal analytics platform, a higher RPO and RTO may be acceptable. The key is to align the DR design with the business requirements, and to document and test the plan. Version note: if you use active-active replication, the RPO can be near zero for the data that is replicated, but conflict resolution becomes a concern. Active-passive is simpler and is the standard for DR; active-active is more complex and is used when both regions must serve traffic simultaneously.
RPO is the maximum tolerable data loss in time; RTO is the maximum tolerable downtime.
RPO is determined by replication lag; RTO by the failover procedure.
Asynchronous replication gives RPO > 0; synchronous or dual-write is needed for RPO=0.
Automated failover is needed for low RTO; manual failover cannot meet a 5-minute RTO.
Run regular failover drills and measure actual RPO and RTO.
Align DR design with business requirements; document and test the plan.
Active-passive is simpler for DR; active-active is more complex but serves both regions.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience