Part 2 · 2 chapters · ~12 min
Memory and CPython Internals
Bytecode and the dis module, the evaluation loop and the 3.11 adaptive interpreter, the object model (PyObject, types, slots), reference counting and the cycle collector, memory allocation with pymalloc, the GIL measured on this machine, free-threaded builds, and C extensions.
4
Bytecode, refcounts and the GIL
code
import dis, sys def fee(a): return max(a * 50 // 10_000, 1_000) dis.dis(fee) # LOAD_FAST a, LOAD_CONST 50, BINARY_OP *, ... RETURN_VALUE x = []; sys.getrefcount(x) # 2 (x plus the temporary argument) # measured on this machine (CPython 3.14.5, GIL enabled), 4 tasks of 3M multiply-adds: # serial 466 ms · 4 threads 455 ms · 4 processes 227 ms
INSIDE CPYTHON
source to bytecode to the evaluation loop, with reference counting and the GIL
swipe the figure sideways, or tap expand for full screen
1/5
compile to bytecode
CPython compiles source to bytecode once and caches it in __pycache__. dis.dis(fn) shows the instructions.
source → bytecode, cached as .pycimport dis; dis.dis(fee)
5
Memory in practice
| topic | what to know |
|---|---|
| object overhead | a small int is ~28 bytes; a dict per instance costs memory: use __slots__ or slotted dataclasses for millions of objects |
| pymalloc | small objects come from arenas; freed memory may not return to the OS, so RSS stays high after a spike |
| leaks | references kept in module-level caches, closures, or cycles with __del__; find them with tracemalloc (Backend Debugging part 3) |
| forked workers | Gunicorn forks share pages copy-on-write until reference counts change them; gc.freeze() before forking reduces copying |
| C extensions | NumPy, Pydantic-core (Rust), orjson do heavy work outside the interpreter and can release the GIL |