10 / 10

Design a worker pool for processing 100,000 jobs with back-pressure and graceful shutdown.

Difficulty: 7/10
worker pool, back-pressure, graceful shutdown

A worker pool uses a buffered jobs channel for back-pressure, N goroutines as workers, a WaitGroup for tracking completion, and context cancellation for graceful shutdown.

Production-grade worker pool
Design notes
  1. 1

    Buffer size controls back-pressure: smaller buffer = tighter back-pressure on producer

  2. 2

    Close the jobs channel to signal workers — ranging over a closed channel exits cleanly

  3. 3

    WaitGroup ensures all in-progress jobs complete before results channel is closed

  4. 4

    For context-aware cancellation, pass ctx to each worker and select on ctx.Done()

  5. 5

    Add a dead-letter queue or retry channel for failed jobs in production

Scenario Questions

0-2 years experience

  1. 1How would you implement a simple worker pool in Go to process 100,000 jobs without overwhelming the system?
  2. 2What happens if you close the job channel while some workers are still processing jobs?
  3. 3How would you signal the workers to stop once all jobs have been dispatched?

2-5 years experience

  1. 1You notice the job queue growing without bound under heavy load; how would you add back‑pressure to the pool?
  2. 2During a deployment the shutdown hangs—what could cause workers not to exit, and how would you fix it?
  3. 3If you need to limit concurrency to 50 workers but also prioritize high‑priority jobs, how would you modify the design?

5-8 years experience

  1. 1Design the pool to handle spikes up to 200,000 jobs while keeping latency low; discuss the trade‑offs between a buffered channel and a semaphore approach.
  2. 2Explain how you would implement graceful shutdown that drains in‑flight jobs, respects a timeout, and reports errors.
  3. 3What monitoring and tuning strategies would you put in place in production to detect deadlocks or resource exhaustion?

8+ years experience

  1. 1If the system must evolve to a distributed worker pool across multiple services, how would you refactor the design while preserving back‑pressure semantics?
  2. 2Discuss the impact of using context cancellation versus a custom shutdown protocol on cross‑team APIs and reliability.
  3. 3What long‑term maintenance concerns arise from coupling job ordering with back‑pressure, and how would you mitigate them?

Follow-up Questions

  • Can you walk me through the shutdown sequence step by step?
  • How would you test that back‑pressure works under load?
  • What metrics would you expose to monitor the pool's health?
Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.