A timeout is an ambiguous outcome, so let the idempotent producer retry internally and make every downstream path tolerate duplicates
The core fact is that a timeout means 'unknown', not 'failed'. The request may have been written and replicated, with only the response lost. So the application must never assume the record is absent. I reason in two layers. Layer one is the producer's internal retries: with enable.idempotence=true, acks=all and a sensible delivery.timeout.ms, the producer retries the same batch with the same sequence number, and the broker deduplicates, so these retries are safe and I do not interfere. Layer two is what happens when the producer finally gives up and the send callback or future reports an error. Now the outcome is still unknown. If the application calls send() again with the same payload, that is a brand new batch with a new sequence number, and the broker cannot deduplicate it. If the first one actually landed, you now have a duplicate.
So my approach is: first, tune delivery.timeout.ms and request.timeout.ms so transient problems are absorbed by internal idempotent retries and the application sees an error only when the cluster has been unhealthy for a long time. Second, treat an application-level resend as at-least-once and make it safe by attaching a stable event ID to every record (generated before the first send and reused on resend) so consumers can deduplicate. Third, for the source-of-truth side, avoid sending directly from request handlers where a lost ack triggers ad hoc retry logic. Write the event to a database outbox in the same transaction as the state change and publish from there, so the retry loop is persistent and replays the same event ID. If the unit of work is read-process-write across Kafka topics, use transactions, which make the write and the offset commit atomic.
Trade-off: a longer delivery.timeout.ms hides more outages from the application but delays failure signals and holds buffer memory. A short one surfaces errors sooner but pushes more ambiguous outcomes onto application logic.
Trade-off: consumer-side deduplication needs a store (a table of seen IDs, or an idempotent upsert keyed by event ID). Naturally idempotent operations, such as setting a value rather than incrementing it, avoid the store entirely.
Common mistake: catching the timeout and calling send() again on the same payload while believing idempotence protects you. It does not, because the new call gets a new sequence number.
Common mistake: generating the event ID at send time instead of at business-event time. A resend then has a different ID and deduplication fails.
Common mistake: treating all exceptions alike. Retriable errors are handled internally; fatal ones such as ProducerFencedException or an out-of-order sequence error generally mean the producer instance should be closed and recreated.
Non-idempotent producers are worse: even the internal retries can duplicate, and with multiple in-flight requests they can reorder.
Version note: idempotence is default from client 3.0, and recovery behavior after a fatal sequence error improved in 2.5 (KIP-360). Confirm the behavior for your client version.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience