Retries with Ordering: Progress vs Order Preservation
This is one of the hardest problems in Kafka consumer design because there is a fundamental tension: retrying a failed record in place preserves ordering but blocks the partition, while moving the failed record aside to keep progress breaks ordering. The right design depends on whether ordering is truly required for all records or only for records with the same key. If ordering is required only per key, you can partition by key and retry per key without affecting other keys. If ordering is required globally within a partition, you have to choose between blocking and breaking order. The first step is to clarify the requirement: most business processes require per-key ordering, not global ordering. For example, all events for a given order must be processed in order, but events for different orders are independent. If that is the case, per-key retries are sufficient and much easier to design.
For per-key ordering, the standard approach is to use a retry topic and a key-based pause. When a record for key K fails, the consumer pauses processing for key K only, while continuing to process other keys. This can be implemented by maintaining an in-memory set of blocked keys and skipping records for those keys until the retry succeeds or the record goes to the DLT. The failed record is produced to a retry topic with its key, and a retry consumer processes it. If it succeeds, the key is unblocked and normal processing resumes. If it exhausts retries, it goes to the DLT and the key is unblocked. This preserves per-key ordering because no other record for key K is processed while the failed record is pending. The trade-off is complexity: you need to track blocked keys, and you need to handle the case where the retry consumer and the main consumer are in the same group or different groups. If you use a single consumer group and the retry topic is consumed by the same group, you need to be careful about rebalancing and offset management.
For global ordering within a partition, the options are more limited. One option is to block the partition and retry in place, accepting that other records are delayed. This is the simplest and preserves ordering, but it means a single poison message can halt the partition indefinitely. To bound this, you set a maximum retry count and then send the record to the DLT, which breaks ordering for that record but unblocks the partition. Another option is to halt the entire consumer group and alert for manual intervention, which is appropriate for high-stakes pipelines where ordering cannot be broken and a human must decide. A third option is to use a single partition with a single consumer and accept the throughput limit. The trade-off is between availability and correctness: blocking preserves order but reduces availability; skipping or DLT preserves availability but breaks order. A common mistake is to assume that global ordering is required when it is not; this leads to over-engineering and unnecessary blocking. Version note: Kafka Streams does not support per-key blocking natively; you would implement it in a custom processor. Spring Kafka's @RetryableTopic does not preserve ordering by default; it moves failed records to retry topics, which means later records for the same key can be processed before the retry succeeds. If ordering is required, you need a custom solution.
Clarify the ordering requirement: per-key vs global. Most processes need per-key only.
Per-key ordering: block only the failed key, continue processing other keys.
Global ordering: block the partition and retry in place, with a max retry count.
Per-key blocking requires coordination between the main consumer and the retry consumer.
Blocking preserves order but reduces availability; DLT preserves availability but breaks order.
Common mistake: assuming global ordering is required when per-key is sufficient.
Spring Kafka @RetryableTopic does not preserve ordering by default.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience