05 / 05

How would you design a background job/task queue system in Python (e.g., using Celery) for handling long-running operations?

Difficulty: 9/10
Task Queues, Celery, Message Brokers, Idempotency

Task queue architecture: broker, workers, idempotency, and observability

A task queue decouples request handling from long-running work. The web request enqueues a job and returns immediately; a worker process picks it up and executes it. The core components are a broker, typically RabbitMQ or Redis, which holds the queue; a worker pool, which consumes jobs and runs them; and a result backend, which stores return values and status if you need to query them. Celery is the most common Python implementation, but RQ, Dramatiq, and cloud queues like SQS are valid alternatives. The hard parts are not the happy path. You must design for idempotency, because most brokers guarantee at-least-once delivery, so a job can run twice. You must handle retries with exponential backoff and a bounded attempt count, and route permanently failing jobs to a dead-letter queue. You must set visibility timeouts and task time limits so a stuck job does not block a worker forever. For observability, you need a dashboard such as Flower, task-level tracing, and metrics for queue depth and processing time. The trade-offs are operational complexity, eventual consistency between the request and the job result, and the cost of running worker infrastructure.

  1. 1

    Broker: RabbitMQ for reliability and routing, Redis for simplicity and speed. Choose based on durability needs.

  2. 2

    Workers: scale horizontally, separate queues by priority or resource type, and set concurrency per worker.

  3. 3

    Idempotency: design tasks so running them twice is safe. Use idempotency keys or check-then-act with a unique constraint.

  4. 4

    Retries: exponential backoff with jitter, bounded attempts, and a dead-letter queue for poison messages.

  5. 5

    Timeouts: set soft and hard time limits so a stuck task cannot block a worker indefinitely.

  6. 6

    Observability: Flower or a custom dashboard, task tracing, queue depth metrics, and alerting on backlog growth.

  7. 7

    Trade-off: Celery is powerful but heavy. For simple needs, RQ or a cloud queue may be easier to operate.

  8. 8

    Common mistake: putting large payloads in the task message. Pass a reference to object storage or a database row instead.

  9. 9

    Version note: Celery 5.x supports Python 3.8+. Broker and result backend compatibility should be verified for your version.

javascript

Scenario Questions

0-2 years experience

  1. 1Why should a web request not wait for a long-running task to finish?
  2. 2What is the role of the broker in a task queue?

2-5 years experience

  1. 1A task runs twice because of a worker crash. How do you make it safe to retry?
  2. 2You need to process urgent tasks before bulk tasks. How do you configure queues and routing?

5-8 years experience

  1. 1Your queue backlog grows without bound during peak hours. How do you diagnose and apply back-pressure?
  2. 2You need task-level tracing across the web request and the worker. How do you propagate context?

8+ years experience

  1. 1Design a task queue architecture for a multi-tenant platform with per-tenant fairness, priority, and isolation, and explain how you test it under failure.
  2. 2Compare Celery, RQ, Dramatiq, and cloud queues on reliability, operational cost, and feature set, and justify a choice for a high-volume service.

Follow-up Questions

  • How do you make a task idempotent when it calls an external payment API?
  • What is the difference between acks_late and acks_early, and when would you use each?
Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.