Stream in chunks or lines with generator pipelines to keep memory constant
The file object itself is already a lazy iterator that yields lines, so the first fix is simply to iterate it instead of calling read() or readlines(). For binary or fixed-size processing, read in bounded chunks with a while loop and yield each chunk. Then compose small generator functions into a pipeline so parsing, filtering, and aggregation each operate on one record at a time. Memory stays proportional to the largest single record plus the pipeline, not the file size. The trade-off is that generators are single-pass: if the algorithm needs two passes, you either re-open the file or spool an intermediate result to disk.
Never call read() or readlines() on a large file; both materialize the whole content.
Wrap file handling in a with block inside the generator so the handle closes even if the consumer abandons the generator early.
For fixed-size binary chunks use a while loop with f.read(size) and break on an empty result.
Trade-off: pure Python per-line generators add interpreter overhead. For very hot paths, batch with itertools.islice or use a compiled parser.
Common mistake: writing sorted(f) or list(generator) at the end of a streaming pipeline, which silently loads everything into memory.
Version note: the walrus operator in while chunk := f.read(size) requires Python 3.8+, so use the explicit break form if you must support 3.7.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience