01 / 05

What is a poison message?

Difficulty: 3/10
Retries, Dead Letter Topic, Poison Message

Poison Messages: Records That Repeatedly Fail Processing

A poison message is a record that repeatedly fails processing no matter how many times it is retried. The failure is deterministic and tied to the content of the record, not to a transient condition like a network blip or a temporarily unavailable downstream service. Examples include a record with a malformed payload that fails deserialization, a record that references a missing foreign key, a record that violates a business rule the consumer enforces, or a record whose schema is incompatible with what the consumer expects. The key characteristic is that retrying the same record with the same code will always produce the same failure. This is what distinguishes a poison message from a transient failure, and it is why infinite retries are the wrong response.

The mechanism that makes poison messages dangerous is head-of-line blocking. In Kafka, a consumer processes records from a partition in order. If the consumer blocks on a poison message, retrying it forever, the offset does not advance, and all subsequent records in that partition are stuck behind it. The consumer appears to be alive but makes no progress, and consumer lag grows without bound. This is especially dangerous in a consumer group because the stuck partition cannot be reassigned to another consumer; the group is stuck until the poison message is handled. The standard remedy is to detect the failure, classify it as permanent, and route the record to a dead-letter topic (DLT) so that the consumer can skip past it and continue processing. This preserves progress at the cost of removing the record from the main flow, which is why the DLT must be monitored and its contents must be replayable after a fix.

A common mistake is to treat every failure as transient and retry indefinitely. This is how a single bad record takes down an entire pipeline. Another mistake is to catch all exceptions and route everything to the DLT without attempting a retry, which discards records that would have succeeded on a second attempt. The correct approach is to classify failures: transient failures (timeouts, connection errors, rate limits) should be retried with backoff; permanent failures (deserialization errors, schema violations, business rule violations) should go to the DLT. A third mistake is to send poison messages to the DLT and never look at them. The DLT is not a graveyard; it is a quarantine that requires monitoring, alerting, and a replay process after the root cause is fixed. The trade-off is between progress and completeness: skipping a poison message keeps the pipeline moving but loses the record unless it is preserved in the DLT and replayed. Version note: Kafka does not have built-in DLT support; you implement it with a producer that writes to a DLT and a consumer that commits past the failed record. Spring Kafka, Kafka Streams, and other frameworks provide DLT abstractions, but the underlying mechanism is the same.

javascript
  1. 1

    A poison message is a record that deterministically fails processing on every retry.

  2. 2

    It causes head-of-line blocking: subsequent records in the partition are stuck behind it.

  3. 3

    Classify failures: transient (retry) vs permanent (DLT).

  4. 4

    Route permanent failures to a dead-letter topic so the consumer can make progress.

  5. 5

    The DLT is a quarantine, not a graveyard; monitor, alert, and replay after fixing the root cause.

  6. 6

    Common mistake: retrying forever and taking down the pipeline for one bad record.

  7. 7

    Kafka has no built-in DLT; frameworks like Spring Kafka provide abstractions.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.