08 / 08

How would you process a very large file that doesn't fit in memory using generators?

Difficulty: 6/10
Generators, Streaming, Lazy Evaluation, Memory

Stream in chunks or lines with generator pipelines to keep memory constant

The file object itself is already a lazy iterator that yields lines, so the first fix is simply to iterate it instead of calling read() or readlines(). For binary or fixed-size processing, read in bounded chunks with a while loop and yield each chunk. Then compose small generator functions into a pipeline so parsing, filtering, and aggregation each operate on one record at a time. Memory stays proportional to the largest single record plus the pipeline, not the file size. The trade-off is that generators are single-pass: if the algorithm needs two passes, you either re-open the file or spool an intermediate result to disk.

  1. 1

    Never call read() or readlines() on a large file; both materialize the whole content.

  2. 2

    Wrap file handling in a with block inside the generator so the handle closes even if the consumer abandons the generator early.

  3. 3

    For fixed-size binary chunks use a while loop with f.read(size) and break on an empty result.

  4. 4

    Trade-off: pure Python per-line generators add interpreter overhead. For very hot paths, batch with itertools.islice or use a compiled parser.

  5. 5

    Common mistake: writing sorted(f) or list(generator) at the end of a streaming pipeline, which silently loads everything into memory.

  6. 6

    Version note: the walrus operator in while chunk := f.read(size) requires Python 3.8+, so use the explicit break form if you must support 3.7.

javascript

Scenario Questions

0-2 years experience

  1. 1What goes wrong with open('big.csv').readlines() on a 20 GB file, and what do you write instead?
  2. 2How do you iterate a text file one line at a time in Python?

2-5 years experience

  1. 1You need the ten largest values from a file that does not fit in memory. How do you combine generators with heapq?
  2. 2Your algorithm needs two passes over a huge file but generators are single-pass. What are your options?

5-8 years experience

  1. 1You chain parsing, filtering, and aggregation stages over a stream. How do you keep memory constant while still getting throughput?
  2. 2You stream a multi-gigabyte object from S3. How do you avoid buffering the whole body, and how do you handle retries mid-stream?

8+ years experience

  1. 1Design a streaming ETL job that is resumable, applies back-pressure, and holds constant memory over a 5 TB input.
  2. 2You need to parallelize CPU-heavy parsing of a stream across cores. How do generators interact with multiprocessing, chunking, and ordering guarantees?

Follow-up Questions

  • How would you take the top 10 values from a file larger than memory without sorting it?
  • What changes if the file is gzipped or read from S3?
Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.