03 / 05

What problem does MirrorMaker 2 solve?

Difficulty: 5/10
Multi-region recovery, RPO/RTO, MirrorMaker

MirrorMaker 2: Cross-Cluster Replication for DR and Migration

MirrorMaker 2 (MM2) solves the problem of replicating data between Kafka clusters, which is needed for disaster recovery, cluster migration, and data aggregation across regions. Kafka's internal replication is designed for a single cluster and does not work across clusters because it requires low-latency, synchronous communication between brokers. MM2 is a separate tool, built on Kafka Connect, that consumes from a source cluster and produces to a target cluster. It replicates topics, partitions, and records, and it also translates consumer offsets so that consumers can resume from the correct position after a failover. MM2 supports active-passive replication (one direction) and active-active replication (both directions), and it can be configured to replicate only selected topics. It solves the problem of how to get data from one cluster to another reliably and at scale, without writing custom replication code. The trade-off is that MM2 is asynchronous, so there is always some replication lag, which determines the RPO. Version note: MM2 replaced MirrorMaker 1, which was a simpler tool with fewer features. MM2 is the current standard and is part of Apache Kafka. It runs on Kafka Connect, so you need a Connect cluster to run it, and it inherits Connect's scaling and fault tolerance.

The mechanism of MM2 is based on Connect source and sink connectors. For each source-target pair, MM2 creates a source connector that reads from the source cluster and a sink connector that writes to the target cluster. It replicates the topic data, and it also replicates internal topics: the checkpoint topic stores the last replicated offset for each partition, the heartbeat topic is used to monitor replication lag, and the offset-syncs topic stores the mapping between source and target offsets. The offset translation is important for failover: when a consumer fails over to the target cluster, it needs to know where to resume. MM2 writes offset mappings so that the consumer can translate its committed offset from the source cluster to the target cluster. Without this, the consumer would have to start from the beginning or from the latest offset, causing duplicates or gaps. MM2 also preserves the topic name by default, but it can rename topics with a prefix to avoid conflicts in active-active setups. The trade-off is between simplicity and control. MM2 is configurable but has many options; a simple setup is easy, but a production DR setup requires careful configuration of topics, offsets, and monitoring. Version note: MM2 is part of Apache Kafka and is actively maintained. It supports both ZooKeeper-based and KRaft-based clusters. Check the version of Kafka and MM2 for feature compatibility.

A common mistake is to assume that MM2 replicates consumer group offsets automatically. It replicates the offset mappings, but the consumer group itself is not replicated; on failover, consumers must be reconfigured to point to the target cluster and to use the translated offsets. Another mistake is to replicate all topics without considering the cost; MM2 can be configured to replicate only selected topics, which reduces cost and complexity. A third mistake is to run MM2 without monitoring replication lag; the heartbeat topic should be monitored, and alerts should fire if the lag exceeds the RPO. The trade-off is between replication completeness and cost. Replicating everything gives the strongest DR but the highest cost; replicating only critical topics is cheaper but requires a clear understanding of what is critical. Version note: MM2 supports replication policies that allow you to include or exclude topics by name or regex, and it supports offset translation for consumers. It also supports 'identity replication' for active-active setups, where topics are replicated with the same name in both directions, which requires careful handling of conflicts. For most DR setups, active-passive with a prefix on the target topics is simpler and safer.

javascript
  1. 1

    MM2 replicates data between Kafka clusters for DR, migration, and aggregation.

  2. 2

    It is built on Kafka Connect and runs as source and sink connectors.

  3. 3

    It replicates topic data and internal topics: checkpoints, heartbeats, offset-syncs.

  4. 4

    Offset translation allows consumers to resume from the correct position after failover.

  5. 5

    It supports active-passive and active-active replication.

  6. 6

    It can replicate selected topics by name or regex to control cost.

  7. 7

    MM2 is asynchronous; replication lag determines RPO.

  8. 8

    MM2 replaced MirrorMaker 1 and is the current standard.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.