Designing a Retry/DLT System for a High-Volume Payment Pipeline
A payment pipeline is a high-stakes environment: records cannot be lost, duplicates can cause double-charges, and failures must be auditable. The design must combine bounded retries, a dead-letter topic, idempotency, and controlled replay. The first step is failure classification. Transient failures include network timeouts, downstream 503s, and rate limits; these should be retried with exponential backoff. Permanent failures include invalid card numbers, insufficient funds, and schema violations; these should go to the DLT immediately or after a small number of confirmatory retries. Ambiguous failures, such as a timeout where the payment may or may not have been processed, are the hardest: they must be handled with idempotency keys so that a retry does not double-charge. The classification should be encoded in the consumer logic, with a policy per exception type.
The retry mechanism should use a series of retry topics with progressively longer delays, for example 1s, 10s, 1m, 10m, and a final DLT. Each retry topic has its own consumer, and the record carries headers for attempt count, original topic, partition, offset, first failure timestamp, and the last exception. Bounded retries are essential: after the maximum number of attempts, the record goes to the DLT, even if the failure is transient. This prevents a single record from blocking the partition indefinitely and makes the failure visible. For a payment pipeline, the retry delays should be tuned to the downstream service's recovery profile: short delays for timeouts, longer delays for rate limits. The retry consumer should be in a separate consumer group from the main consumer so that retries do not block main processing, and it should use the same idempotency key as the main consumer so that a retry that succeeds after the original actually succeeded does not double-charge.
Idempotency is the linchpin. Every payment record must carry a unique idempotency key, and the downstream payment processor must enforce it. If the processor does not support idempotency keys, you need an inbox table or a deduplication store in your own system. The DLT should be a compacted or retention-bounded topic with enough context to diagnose and replay: the original record, the headers, and the exception. Replay must be controlled: you do not want to blindly replay the DLT into the main topic, because that can cause duplicates or re-trigger failures. Instead, provide a replay tool that reads from the DLT, allows filtering and transformation, and produces to a replay topic that is consumed by the same consumer logic but with a flag indicating it is a replay. The consumer should be idempotent so that replay is safe. Auditability requires that every state transition (retry, DLT, replay) is logged and traceable. Observability requires metrics on retry rates, DLT volume, replay volume, and the age of the oldest record in each retry topic and the DLT. Alerts should fire when DLT volume exceeds a threshold or when the oldest record is too old. The trade-off is between complexity and safety: more retry topics and more idempotency infrastructure add operational overhead, but for payments, the cost of a double-charge or a lost payment is far higher.
Classify failures: transient (retry with backoff), permanent (DLT), ambiguous (idempotency).
Use retry topics with progressive delays: 1s, 10s, 1m, 10m, then DLT.
Bounded retries: after max attempts, go to DLT even for transient failures.
Idempotency key per payment; enforce at the processor or with an inbox table.
Retry consumer in a separate group; same idempotency key as main consumer.
Controlled replay: filter and transform DLT records, produce to a replay topic.
Auditability and observability: log every state transition, alert on DLT volume and age.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience