Challenging the Claim of Exactly-Once Business Processing Across Kafka, a Database, and an External Payment Provider
The claim is incorrect, and the challenge is to separate Kafka's exactly-once semantics from distributed transactions. Kafka's exactly-once semantics (EOS) guarantee that a consume-transform-produce pipeline within Kafka can read from a topic, process a record, and write to another topic exactly once, even in the presence of failures. This is achieved with idempotent producers and transactions. But Kafka transactions only cover Kafka reads and writes. They do not cover a database write or an external payment provider call. If the team writes to a database and calls a payment provider inside the processing logic, those side effects are not part of the Kafka transaction. If the Kafka transaction commits but the database write fails, or the database write succeeds but the Kafka transaction aborts, the system is inconsistent. The only way to achieve exactly-once across Kafka, a database, and an external payment provider is to use a distributed transaction (e.g., XA) or a saga with compensating actions, and even then, exactly-once is not guaranteed for the external provider unless it supports idempotency. The correct claim is: Kafka provides exactly-once within Kafka; for external systems, you get at-least-once plus idempotency, which is effectively-once for business purposes if the idempotency is implemented correctly.
The mechanism that limits Kafka's EOS is the transaction boundary. A Kafka transaction can include multiple topic partitions and the consumer offsets, but it cannot include an external database or an API call. The transaction coordinator manages the two-phase commit for Kafka resources only. If the processing logic writes to a database, that write is not part of the Kafka transaction, so it can succeed or fail independently. If the processing logic calls a payment provider, the call is not part of the transaction, and the provider may or may not support idempotency. To achieve effectively-once across all three, you need: (1) Kafka transactions for the Kafka-to-Kafka path, (2) the outbox pattern or idempotency for the database, and (3) idempotency keys for the payment provider. If the payment provider does not support idempotency keys, you cannot guarantee exactly-once; you can only guarantee at-least-once with reconciliation. The trade-off is between the complexity of distributed transactions and the practicality of idempotency plus reconciliation. For most systems, idempotency plus reconciliation is the right choice because distributed transactions are slow, complex, and not supported by all systems. Version note: some payment providers support idempotency keys (e.g., Stripe's Idempotency-Key header), which makes effectively-once achievable. If the provider does not, you must build a reconciliation process to detect and correct duplicates. Always check the provider's capabilities before claiming exactly-once.
A common mistake is to assume that Kafka transactions cover the entire processing logic, including database writes and API calls. They do not. Another mistake is to assume that the external provider is idempotent without verifying. A third mistake is to skip reconciliation, which is the safety net for any system that cannot guarantee exactly-once. The trade-off is between the cost of implementing idempotency and reconciliation and the risk of duplicates or loss. For a financial system, the cost of a duplicate payment is high, so idempotency and reconciliation are essential. The correct approach is to: use Kafka transactions for the Kafka path, use the outbox pattern for the database, use idempotency keys for the payment provider, and implement reconciliation to detect and correct any discrepancies. The claim of exactly-once should be replaced with effectively-once, which is the realistic guarantee for a distributed system with external side effects. Version note: the terminology matters. Exactly-once is a strong guarantee that is hard to achieve across systems. Effectively-once means the system behaves as if it were exactly-once, using idempotency and reconciliation. Most production systems achieve effectively-once, not exactly-once. Be precise in the claim.
Kafka EOS covers Kafka reads and writes, not external systems.
Database writes and payment provider calls are outside the Kafka transaction.
Exactly-once across all three requires a distributed transaction or idempotency plus reconciliation.
Most payment providers support idempotency keys; verify before claiming exactly-once.
Reconciliation is the safety net for any system that cannot guarantee exactly-once.
The correct claim is effectively-once, not exactly-once, for external side effects.
Use Kafka transactions for the Kafka path, outbox for the database, idempotency keys for the provider.
Be precise in the claim; exactly-once is a strong guarantee that is hard to achieve across systems.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience