Testing Consumer Behavior When the Broker Becomes Unavailable
Testing broker unavailability is about verifying that the consumer handles the failure gracefully: it should retry the connection, not lose records, not commit offsets prematurely, and not produce duplicate side effects. The test must simulate the broker going down while the consumer is processing a batch, then bring the broker back and verify that the consumer resumes correctly. With Testcontainers, you can pause or stop the Kafka container to simulate unavailability, though stopping and restarting the container is heavier than pausing. A cleaner approach is to use a Toxiproxy container between the consumer and the broker to introduce network failures, latency, or connection resets without stopping the broker. This gives you fine-grained control over the failure and makes the test faster and more deterministic. The test should assert on three things: the consumer did not commit the offset for records that were not processed, the consumer reconnected after the failure, and the records were eventually processed exactly once (if idempotent) or at least once (if not).
The mechanism for testing this is to control the failure injection and the consumer's configuration. The consumer should have enable.auto.commit=false so that offsets are committed manually after processing. During the failure, the consumer's poll() will throw or return no records; the consumer should catch the exception, log it, and retry with backoff. When the broker comes back, the consumer should rejoin the group and resume from the last committed offset. If the failure happened after processing but before commit, the records will be re-delivered, which is why the consumer must be idempotent. The test should verify that the side effects (e.g., database writes) are applied only once even though the records were delivered twice. This is where the test becomes valuable: it proves that the idempotency mechanism works under failure. The trade-off is between test complexity and confidence. Simulating broker unavailability is more complex than a happy-path test, but it is exactly the scenario that causes production incidents. Version note: Toxiproxy is a common tool for network fault injection and has a Testcontainers module. Kafka's client libraries have retry and reconnect logic built in, but the behavior depends on configs like reconnect.backoff.ms, retry.backoff.ms, and session.timeout.ms; the test should exercise the actual configs used in production.
A common mistake is to test only the happy path and assume that the client library handles failures. It does handle many failures, but the application code around it (commit logic, idempotency, error handling) is where bugs live. Another mistake is to use Thread.sleep to simulate a failure; this does not actually break the connection and does not test the reconnection logic. A third mistake is to assert that the consumer processed the records exactly once without verifying that the failure actually occurred; if the failure injection did not work, the test passes for the wrong reason. The trade-off is between realism and speed. Stopping the broker is realistic but slow and can leave the container in a bad state. Toxiproxy is faster and more controllable but adds a component to the test. For most teams, Toxiproxy is the better choice because it makes failure tests deterministic. The test should also verify that the consumer does not enter a tight retry loop that hammers the broker; the backoff should be observable in the test logs or metrics.
Simulate broker unavailability with Toxiproxy or by stopping the container.
Use enable.auto.commit=false and manual commits to control offset semantics.
Assert that unprocessed records are not committed and are re-delivered after recovery.
Verify idempotency: side effects applied exactly once despite re-delivery.
Avoid Thread.sleep as a failure simulation; use real network fault injection.
Check that the consumer uses backoff and does not hammer the broker.
Toxiproxy is faster and more deterministic than stopping the broker container.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience