Consumers pull: poll() drives fetch requests to partition leaders, records are buffered locally, and the application processes them at its own pace
Kafka uses a pull model. The broker never pushes records to a consumer; the consumer asks for them. Inside the client, a fetcher sends fetch requests to the leaders of the partitions the consumer owns, receives batches of records, and holds them in an in-memory buffer. When the application calls poll(), the client returns up to max.poll.records records from that buffer, and it usually has already sent the next fetch so data is ready by the time the application finishes processing. Processing is entirely the application's job, and so is the position: the consumer tracks an offset per partition, and committing that offset to the __consumer_offsets topic is a separate, explicit step from reading.
Why pull matters: the consumer controls its own rate, so a slow consumer builds lag instead of being overwhelmed, and a new or recovering consumer can read at full speed from any offset. poll() is also the heartbeat of the consumer's liveness: it triggers group join and rebalance handling, and the time between polls is monitored. A common misconception is that poll() blocks until it has something to return for as long as you like. The timeout you pass is only the maximum wait, and the call can return empty.
Trade-off: pull gives backpressure and replay but means the consumer must poll regularly even when idle, and polling too slowly has consequences (see the max.poll.interval.ms question).
Trade-off: auto-commit is convenient but commits on a timer inside poll(), so a crash after a commit but before processing finishes can lose work. I turn it off and commit after processing for anything that matters.
Common mistake: assuming the consumer is thread-safe. KafkaConsumer is not; use one consumer per thread, or hand work off to other threads and keep all consumer calls on the polling thread (wakeup() is the only exception).
Common mistake: confusing reading with committing. A record you have received is not 'consumed' from the broker's point of view until you commit an offset, and the broker deletes nothing either way.
Common mistake: doing long blocking work in the poll loop without understanding the poll interval limit, which can get the consumer kicked out of its group.
Version note: use poll(Duration); the old poll(long) was deprecated long ago and is gone in recent major versions. Kafka 4.0 also makes the KIP-848 group protocol generally available, which moves group coordination logic to the broker. The poll and fetch model itself is unchanged, but rebalance-related configuration names differ, so check the version you run.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience