Interpretation, dynamic typing, and boxing add per-operation overhead; mitigate with C extensions, vectorization, or compilation
CPython compiles Python source to bytecode, not machine code, and executes that bytecode in a loop written in C. Each bytecode is a dynamic dispatch: the interpreter must look up the operation on a type object, check argument types, and allocate a new result object. Numbers are heap-allocated PyLong and PyFloat objects, so a simple addition is a function call plus an allocation rather than a machine add. Attribute lookups go through the MRO and the descriptor protocol. None of these individual costs is large, but they apply per operation, and a tight loop executes millions of them. C++ pays those costs once at compile time and emits direct machine instructions with registers and stack values. That is the whole reason for the gap: not that Python is poorly implemented, but that it trades runtime dynamism for compile-time specialization.
Move the hot loop out of Python: use NumPy, Cython, or a C extension so the loop runs in compiled code.
Use built-ins and stdlib functions where possible. They are implemented in C and avoid per-item bytecode.
Reduce allocation: preallocate lists, avoid building intermediate objects, and reuse buffers.
Cache attribute lookups and method references in hot paths, since lookup is not free.
Consider PyPy for long-running pure Python CPU work, or Numba for numeric kernels, but verify library compatibility.
Trade-off: C extensions reduce portability and add build complexity. Vectorization requires the problem to fit an array model.
Common mistake: micro-optimizing Python syntax while leaving an O(n^2) algorithm in place. Fix the algorithm first.
Version note: Python 3.11 introduced an adaptive specializing interpreter and 3.13 continued this work, narrowing the gap for some workloads but not eliminating it.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience