02 / 05

How would you implement delayed retries for transient failures?

Difficulty: 4/10
Retries, Dead Letter Topic, Poison Message

Delayed Retries with Retry Topics and Backoff

Delayed retries are about retrying a failed record after a waiting period instead of immediately. Immediate retries in a tight loop are almost always wrong: they hammer the failing downstream service, they block the consumer from making progress on other records, and they often fail again because the transient condition has not resolved. The standard Kafka pattern is to use a series of retry topics, each with a progressively longer delay. When a record fails, the consumer produces it to retry-topic-1 with a header indicating the next attempt time, and commits the offset on the main topic. A separate consumer (or the same consumer with a different subscription) reads from the retry topic, waits until the next attempt time, and re-processes the record. If it fails again, it produces to retry-topic-2 with a longer delay, and so on, until it either succeeds or exhausts the retry attempts and goes to the DLT. This pattern keeps the main consumer moving and gives the downstream service time to recover.

The mechanism for the delay can be implemented in several ways. One approach is to use a consumer that pauses between polls, but this is inefficient and ties up a consumer thread. A better approach is to use a scheduler that checks the next-attempt header and only processes records that are due. Some frameworks, like Spring Kafka, provide a @RetryableTopic annotation that creates retry topics automatically and uses a listener container that respects the delay. Another approach is to use a single retry topic and a consumer that partitions by the next-attempt time, but Kafka does not support time-based consumption natively, so you would need to poll and filter. The most common production pattern is multiple retry topics with fixed delays (e.g., 1s, 10s, 1m, 10m) and a final DLT. This is simple to operate and easy to reason about. The trade-off is between the number of retry topics and the granularity of the backoff. More topics give finer backoff but more infrastructure to manage. Fewer topics are simpler but may not fit the failure profile.

A common mistake is to retry in the same consumer thread with Thread.sleep, which blocks the entire partition and causes head-of-line blocking for unrelated records. Another mistake is to retry forever with increasing backoff; at some point, you need to give up and send the record to the DLT, otherwise a permanent failure will loop forever. A third mistake is to lose the original offset and headers when moving to the retry topic; without them, you cannot trace the record back to its source or replay it correctly. The trade-off is between latency and load: longer delays reduce load on the downstream service but increase the time to eventual success. For a payment pipeline, you might use short delays (seconds) for retryable errors like timeouts and longer delays (minutes) for rate limits. Version note: Spring Kafka's @RetryableTopic and the non-blocking retry pattern were introduced in Spring Kafka 2.7+; if you are on an older version, you need to implement the pattern manually. Kafka itself does not provide retry topics; they are a pattern, not a feature.

javascript
  1. 1

    Use retry topics with progressively longer delays instead of tight retry loops.

  2. 2

    Produce the failed record to the next retry topic and commit the offset on the main topic.

  3. 3

    Carry headers: attempt count, original topic, original offset, first failure time.

  4. 4

    After exhausting retries, send to the DLT, not back to the main topic.

  5. 5

    Avoid Thread.sleep in the consumer thread; it blocks the partition.

  6. 6

    Spring Kafka's @RetryableTopic automates the pattern (2.7+).

  7. 7

    Trade-off: longer delays reduce downstream load but increase time to success.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.