05 / 05

Design multi-region Kafka for a globally distributed financial platform.

Difficulty: 9/10
Multi-region recovery, RPO/RTO, MirrorMaker

Multi-Region Kafka for a Globally Distributed Financial Platform

A globally distributed financial platform has three requirements that drive the multi-region design: low latency for users in different regions, high availability with no data loss, and regulatory compliance with data residency. The first decision is the topology: active-passive, active-active, or a hybrid. Active-passive is simpler and is the standard for disaster recovery: one region handles all traffic, and a standby region is kept in sync for failover. Active-active means both regions handle traffic simultaneously, which gives lower latency and better utilization but requires conflict resolution and careful ordering. For a financial platform, active-passive is often the right choice for the core ledger, because it avoids the complexity of conflict resolution and preserves a single source of truth. Active-active can be used for read-only or regional services, such as displaying balances or processing regional payments, with the core ledger remaining active-passive. The second decision is replication: use MirrorMaker 2 for asynchronous replication between regions, and monitor replication lag to ensure it meets the RPO. For RPO=0, synchronous replication is needed, which can be achieved with a dual-write pattern or a consensus-based system, but this adds latency and complexity. The trade-off is between latency, consistency, and availability. A financial platform usually prioritizes consistency and durability over latency, so active-passive with asynchronous replication and a well-tested failover is a common design.

The mechanism for the design has several layers. At the data layer, topics are replicated with MM2 from the primary region to the standby. The replication includes the topic data and the internal topics (checkpoints, heartbeats, offset-syncs) that enable offset translation. At the application layer, producers and consumers are configured to use a regional endpoint that resolves to the primary cluster; on failover, the endpoint is updated to point to the standby. At the schema layer, the schema registry must also be replicated or made available in both regions, because producers and consumers need to resolve schemas. At the governance layer, data residency requirements may dictate that certain data cannot leave a region; in that case, the replication must be selective, and some topics may not be replicated. At the operational layer, failover procedures must be documented and tested, and the team must run regular drills to measure the actual RPO and RTO. The trade-off is between complexity and correctness. A multi-region financial platform is complex to build and operate, but the cost of data loss or downtime is far higher. Version note: MM2 supports offset translation and can replicate the schema registry if it is configured as a Connect cluster. Some managed services provide multi-region replication out of the box; if you use one, check its guarantees and test the failover.

A common mistake is to choose active-active for the core ledger without a plan for conflict resolution. Two regions accepting writes to the same account can cause overdrafts or double-spends. Another mistake is to ignore data residency: replicating all data to all regions may violate regulations. A third mistake is to assume that asynchronous replication gives zero RPO; it does not, and the business must accept the gap or pay for synchronous replication. The trade-off is between latency and consistency. Active-active gives lower latency but weaker consistency; active-passive gives stronger consistency but higher latency for users in the standby region. For a financial platform, the core ledger should be active-passive with a clear RPO and RTO, and regional services can be active-active for reads. A good design is a hybrid: the ledger is active-passive, while the API layer is active-active and routes writes to the primary region. Version note: if you need RPO=0 across regions, you cannot use Kafka's internal replication; you need a synchronous dual-write or a consensus layer. Some financial platforms use a separate system (e.g., a distributed database) for the ledger and use Kafka for event distribution, which decouples the consistency requirements. Always align the design with the business requirements and the regulatory constraints.

javascript
  1. 1

    Choose active-passive for the core ledger; active-active for reads and regional services.

  2. 2

    Use MirrorMaker 2 for asynchronous cross-region replication and monitor lag for RPO.

  3. 3

    Replicate the schema registry or make it available in both regions.

  4. 4

    Respect data residency; replicate only the topics that are allowed to cross regions.

  5. 5

    Use DNS or a configuration service to redirect producers and consumers on failover.

  6. 6

    Document and test the failover procedure; measure RPO and RTO with drills.

  7. 7

    For RPO=0, use synchronous dual-write or a consensus layer, not Kafka's internal replication.

  8. 8

    Align the design with business requirements and regulatory constraints.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.