Threads release the GIL around blocking I/O, so waiting overlaps; CPU work never releases it
The GIL is released automatically before blocking system calls such as socket reads and writes, file reads, and sleeps, and it is reacquired when the call returns. For I/O-bound work this means a thread that is waiting on the network does not block other threads from running Python code, so N threads can have N outstanding requests in flight and total wall time drops. For CPU-bound work there is no release point: a tight Python loop holds the GIL and only yields every few milliseconds due to the switch interval, which adds context-switch overhead without giving real parallelism. The result is that threading helps latency-bound and IO-bound workloads, but for CPU-bound workloads the only way to get true parallelism in CPython is multiple processes or GIL-releasing C extensions.
I/O-bound: threads overlap waiting; this is the classic concurrency win.
CPU-bound: threads serialize on the GIL and add overhead; measured speedup is often below 1x.
GIL-releasing C extensions (numpy, hashlib, zlib, database drivers) can run in real parallel even inside threads.
sys.setswitchinterval can tune how often the interpreter checks for a thread switch, but it does not create CPU parallelism.
Trade-off: threads share memory so coordination is cheap but races are possible. Processes avoid races but need serialization and IPC.
Common mistake: benchmarking a threaded CPU loop on a single core and concluding the GIL is not a problem.
Version note: the GIL is released around blocking IO in all Python 3 versions; the free-threaded build in 3.13+ changes the CPU-bound picture but is opt-in.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience