Snapshot with tracemalloc, inspect gc.get_objects(), find the growth, then break the retention
Python's GC handles cycles, so a growing RSS means something is keeping objects alive that you did not intend. The classic causes are module-level caches with no bound, event handlers or callbacks registered and never removed, closures capturing large objects, thread locals, and references held by C extensions. The systematic approach is to measure first: use tracemalloc to take snapshots at intervals and compare the top allocation sites, which shows which lines are responsible for growth. Then use gc.get_objects() and gc.get_referrers() to find who is holding a growing object. For long-lived services, a heap snapshot diff is the fastest path. Once you find the retention path, the fix is usually bounding the cache with an LRU, using weak references, or unregistering callbacks on teardown.
tracemalloc.start(), take two snapshots minutes apart, then snapshot.compare_to(previous, 'lineno') gives the top growth lines.
gc.get_objects() lets you count how many instances of a type exist; a steadily increasing count is a strong signal.
gc.get_referrers(obj) shows who holds a reference, which usually reveals the cache or registry keeping it alive.
objgraph and pympler give more readable object graphs and diff views.
weakref.WeakValueDictionary or WeakSet is the standard fix for caches that must not keep objects alive.
Common mistake: assuming gc.collect() will fix a leak. It only breaks cycles; it does not release objects still referenced.
Version note: tracemalloc is available since 3.4. gc.freeze() in 3.7 helps fork-based servers by moving startup objects out of GC scans.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience