Offload CPU work to a process pool via run_in_executor or asyncio.to_thread
The event loop is single-threaded, so any CPU-bound code inside an async function stalls every other task. The fix is to hand that work to a worker that does not run on the loop thread. For CPU-bound work, use a ProcessPoolExecutor and await loop.run_in_executor(pool, func, *args) or wrap it in an async helper. A ThreadPoolExecutor helps only for blocking IO or for C code that releases the GIL; for pure Python CPU work it will just contend on the GIL. The important design points are: bound the pool size to the available cores, avoid passing huge objects across process boundaries because they are pickled, and apply back-pressure so the queue of submitted work does not grow without bound. In 3.9+ asyncio.to_thread is a convenient shortcut for the blocking-IO case.
ProcessPoolExecutor for CPU-bound work; ThreadPoolExecutor only for blocking IO or GIL-releasing C calls.
loop.run_in_executor is awaitable and integrates cleanly with gather, timeouts, and cancellation.
Bound the pool size to os.cpu_count() or a configured limit; the default can overwhelm the machine.
Pickle cost is real. Chunk data and pass file paths or shared memory, not multi-megabyte objects.
Apply back-pressure with an asyncio.Semaphore or a bounded queue so you do not submit faster than the pool can drain.
Common mistake: passing an open file object, a socket, or a database connection into a process. They are not picklable or meaningful in the child.
Common mistake: creating a new pool per request. Pool startup is expensive; create one pool for the service lifetime.
Version note: asyncio.to_thread was added in 3.9 as a convenience wrapper. ProcessPoolExecutor has been in concurrent.futures since 3.2.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience