06 / 12

What causes a reference cycle, and how does Python's garbage collector handle it?

Difficulty: 9/10
Variables, Data Types & Memory Model, Garbage Collection, Reference Cycles and Weak References

Reference cycles and generational cyclic GC

A reference cycle exists when objects reference each other directly or indirectly (a list containing itself, parent and child pointing at each other, a closure capturing the object that owns it). Each object keeps a non-zero reference count from the others, so pure reference counting never reaches zero even after the program loses all external references. CPython therefore adds a supplementary cycle collector, exposed through the gc module.

How it works: the collector only tracks container objects (lists, dicts, sets, class instances and so on); ints and strings cannot form cycles and are skipped. For a generation being collected it copies each tracked object's refcount, subtracts the references that come from other tracked objects in the same set, and any object left with a count above zero is reachable from outside. Everything reachable from those roots is kept; the rest is unreachable garbage and gets finalized and freed. Objects are placed into three generations; new objects start young and survive into older generations, which are collected less often, based on the generational hypothesis that most objects die young. Default thresholds are visible with gc.get_threshold() (classically 700, 10, 10).

javascript

Trade-offs and alternatives: cycle collection is not free; collections cause latency pauses that grow with the number of tracked objects. Better than relying on GC is not creating cycles: use weakref for back-pointers (child to parent, observer lists, caches). Teams with strict latency or memory-sharing needs sometimes call gc.disable() and gc.freeze() after startup (as Instagram famously described), but then you take responsibility for avoiding cycles and for occasional manual collection. gc.disable() as a general 'speed-up' is a common mistake because leaks follow.

Version-dependent details: since 3.4 (PEP 442) objects with del in cycles can be collected safely, whereas in Python 2 they were uncollectable. The incremental GC was developed during the 3.13 cycle, reverted before the 3.13 release, and later released in 3.14, so check the release notes of your target version. In the free-threaded build the collector works differently. Also, a gc.collect() count is a diagnostic, not a guarantee of how much memory returned to the OS.

Scenario Questions

0-2 years experience

  1. 1Given a = []; a.append(a); del a - is the memory freed immediately? Why or why not?
  2. 2What does gc.collect() return, and in what situations might you call it manually?

2-5 years experience

  1. 1A tree's children keep a reference back to their parent. How do you avoid creating cycles with weakref, and what trade-offs does that introduce?
  2. 2A service shows irregular latency spikes that you suspect come from garbage collection. How would you confirm it and what mitigations would you consider?

5-8 years experience

  1. 1Explain the three generations and thresholds, and how you would tune them for a batch job that allocates millions of long-lived objects.
  2. 2When is gc.disable() combined with gc.freeze() a reasonable production choice, and what risks does it introduce?

8+ years experience

  1. 1Walk through the algorithm CPython uses to find unreachable cycles, including why only containers are tracked and how finalizers and object resurrection are handled (PEP 442).
  2. 2Compare CPython's refcounting-plus-cycle-GC with a tracing GC (PyPy or the JVM) on latency, throughput and determinism, and explain what this means for writing portable library code.

Follow-up Questions

  • Why does the cycle collector only track container objects?
  • What problem did PEP 442 solve for objects with __del__ in cycles?
Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.