Replay with an isolated group.id or by resetting only the target group's offsets, since offsets are per group and other groups are unaffected
Because committed offsets are stored per group, replay is naturally isolated: changing one group's position never touches another group's. That gives me two clean options. Option one is a dedicated replay group: start a separate application or job with a new group.id and seek it to the chosen starting point. It cannot disturb the production group, it can be throttled and stopped independently, and the production group keeps its position. Option two is resetting the production group itself with kafka-consumer-groups.sh --reset-offsets, which is appropriate when that same service needs to reprocess (for example after a bug fix), but the group must have no active members while the reset executes, and the reset rewinds only that group.
The starting point can be the earliest offset, a specific offset, a shift relative to the current offset, or a timestamp. Timestamp-based resets resolve to the earliest offset whose record timestamp is at or after the time, via offsetsForTimes. The checks that matter more than the mechanics are these. Is the data still there? Replay is limited by retention, and compacted topics only retain the latest value per key. What happens downstream? A replay re-triggers every side effect, so handlers must be idempotent or write to a separate sink. Is the cluster protected? Reading old data comes from disk rather than page cache, which can hurt broker latency for other clients, so apply a consumer fetch quota to the replay client and run it off-peak.
Trade-off: a separate replay group is safest and cleanly separable, but if the same service must reprocess, you eventually reset the real group, which requires stopping it and coordinating the cutover.
Trade-off: manual assign() with seek avoids group coordination and rebalances during replay, but gives no automatic failover. For a one-off job that is acceptable.
Common mistake: resetting offsets while consumers are running. The tool refuses for active groups, and forcing the situation by restarting consumers mid-reset causes confusing results.
Common mistake: forgetting that a reset rewinds the group for the specified topics only. Other topics the group subscribes to are untouched unless you pass --all-topics.
Common mistake: replaying into production side effects (emails, payments, webhooks). Replay to a sink that is safe, or make every handler idempotent, and dry-run before executing.
Common mistake: replaying data that no longer exists. Verify retention, log start offsets and compaction before promising a replay window. Tiered storage can extend history, but check your version and configuration.
Operational note: record the replay window, the resulting offsets and the owner. After finishing, delete the temporary group so it does not linger with stale offsets and lag alerts.
Version note: reset tooling and flags are stable across modern versions, but behavior of group state and describe output differs slightly under the KIP-848 protocol. Verify that the group is truly empty before executing a reset.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience