nextRound
TechnologiesCoding ProblemsBookmarksLearning PathsLogin
nextRound
TechnologiesCoding ProblemsBookmarksLearning PathsLogin
nextRound

AI-powered interview preparation platform. Practice with curated questions, mock interviews, and personalized learning paths to crack your dream tech interview.

Quick Links

  • Technologies
  • Mock Interviews
  • Saved Questions
  • Pricing

Company

  • About Us
  • Contact Us

Legal

  • Privacy Policy
  • Terms of Use

© 2026 nextRound. All rights reserved.

Questions
1 of 5
1How would you design a rate limiter for an API built in Python?
2How would you architect a Python web service to handle 10x traffic growth?
3How would you design a plugin system in Python that allows third parties to extend your application's functionality?
4What tradeoffs would you consider when choosing between a synchronous framework (e.g., Flask/Django) and an asynchronous one (e.g., FastAPI) for a new service?
5How would you design a background job/task queue system in Python (e.g., using Celery) for handling long-running operations?
PythonPython
Basics
Control Flow and Functions
Data Structures
Comprehensions & Functional Programming
Iterators, Generators & Decorators
Object-Oriented Programming
Exception Handling & Debugging
Concurrency & Parallelism
Performance & Optimization
Testing
Security
Modules, Packaging & Environment
Type Hinting & Modern Python
System Design & Architecture with Python
Best Practices & Design Patterns
Edge Cases & Tricky Interview Questions
01 / 05

How would you design a rate limiter for an API built in Python?

Difficulty: 9/10
Rate Limiting, Token Bucket, Sliding Window, Distributed Systems

Rate limiter design: algorithm choice, shared storage, and atomic operations

A rate limiter controls how many requests a client can make in a time window. The first decision is the algorithm. Token bucket and leaky bucket allow bursts up to a bucket size and then smooth to a refill rate, which matches most API use cases. Fixed window counters are simple but have a boundary problem: a client can send a full burst at the end of one window and another at the start of the next, doubling the effective rate. Sliding window log stores timestamps and gives exact counts but uses memory proportional to the number of requests. Sliding window counter approximates by weighting the previous window and is a good middle ground. The second decision is storage. A single-process limiter can use an in-memory dict, but any multi-instance service needs a shared store, typically Redis, with atomic operations. Doing read-modify-write in Python is a race condition under concurrency; you need a Lua script or an atomic command like INCR with EXPIRE, or Redis sorted sets with ZADD and ZREMRANGEBYSCORE for sliding windows. The third decision is scope and failure mode: per user, per API key, per IP, or global. Decide whether to fail open or closed when Redis is unavailable, and make that explicit in the design.

  1. 1

    Token bucket: O(1) per request, supports bursts, easy to reason about. Store tokens and last refill time per key.

  2. 2

    Sliding window log: exact, no boundary problem, but memory grows with request rate. Use sorted sets in Redis.

  3. 3

    Fixed window: simplest, but burst at boundaries can double the limit. Only acceptable for coarse limits.

  4. 4

    Distributed: use Redis with atomic Lua scripts or built-in atomic ops. Never do get-then-set from Python.

  5. 5

    Trade-off: exactness versus memory and latency. Token bucket is usually the best default for APIs.

  6. 6

    Common mistake: implementing the limiter in the web process and assuming it works across pods. It does not.

  7. 7

    Common mistake: using time.time() in multiple processes without a shared clock; use Redis server time or monotonic clocks where possible.

  8. 8

    Version note: Redis 7+ supports functions; Lua scripts have been supported for a long time. Use whichever your infrastructure supports.

javascript

Scenario Questions

0-2 years experience

  1. 1Why is a fixed window counter vulnerable to burst traffic at window boundaries?
  2. 2How do you implement a simple in-memory rate limiter for a single process?

2-5 years experience

  1. 1You have three web pods behind a load balancer. How do you make the rate limiter consistent across them?
  2. 2A client sends 1000 requests per second and you must limit to 100. Which algorithm do you choose and why?

5-8 years experience

  1. 1You need per-endpoint rate limits with different capacities and refill rates. How do you design the key space and configuration?
  2. 2Your rate limiter adds 5ms latency per request. How do you reduce that without losing correctness?

8+ years experience

  1. 1Design a rate limiting layer that supports per-tenant quotas, burst allowances, and graceful degradation when the shared store is unavailable.
  2. 2Explain how to test a distributed rate limiter for correctness under clock skew, network partitions, and concurrent bursts.

Follow-up Questions

  • How would you rate limit per user and per IP simultaneously without doubling the storage cost?
  • What happens when Redis is down and you fail open? How do you protect the backend?
Sharethis question

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.