Built-ins run in C and NumPy uses contiguous buffers plus vectorized kernels
A pure Python loop pays interpreter overhead for every iteration: bytecode dispatch, bounds checks, type lookups, and result allocation. Functions like sum, min, max, sorted, and map are implemented in C, so the loop runs inside the interpreter without per-iteration bytecode dispatch. NumPy goes further: it stores homogeneous data in a contiguous buffer and applies operations as compiled loops over that buffer, so the per-element cost drops from interpreter overhead to a few machine instructions, and the work is cache-friendly. The result is typically one to two orders of magnitude faster for numeric workloads. The trade-off is that you pay an up-front allocation and that operations that do not fit the vector model, such as dependent branching per element, do not vectorize well.
Prefer sum, min, max, sorted, any, all, and itertools over hand-written loops when the logic matches.
NumPy operations are vectorized: express the whole computation as array expressions, not element-by-element Python loops.
Avoid mixing Python scalars and arrays in a hot loop; each conversion allocates and defeats the vectorization.
Use array-level operations like np.where, np.einsum, and broadcasting instead of Python branches inside a loop.
Trade-off: NumPy uses more memory per array than a list of small ints in some cases, and does not support arbitrary Python objects efficiently.
Common mistake: writing a Python for loop over a NumPy array, which is slower than the equivalent list comprehension because each element is boxed.
Version note: NumPy 2.0 changed some promotion rules and removed some aliases. Verify behavior when upgrading.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience