Part 2 · 2 chapters · ~12 min

Memory and CPython Internals

Bytecode and the dis module, the evaluation loop and the 3.11 adaptive interpreter, the object model (PyObject, types, slots), reference counting and the cycle collector, memory allocation with pymalloc, the GIL measured on this machine, free-threaded builds, and C extensions.

4

Bytecode, refcounts and the GIL

code
import dis, sys
def fee(a): return max(a * 50 // 10_000, 1_000)
dis.dis(fee)                 # LOAD_FAST a, LOAD_CONST 50, BINARY_OP *, ... RETURN_VALUE
x = []; sys.getrefcount(x)   # 2 (x plus the temporary argument)

# measured on this machine (CPython 3.14.5, GIL enabled), 4 tasks of 3M multiply-adds:
#   serial 466 ms · 4 threads 455 ms · 4 processes 227 ms
INSIDE CPYTHON
source to bytecode to the evaluation loop, with reference counting and the GIL
.py sourcecompilerAST → bytecode__pycache__/*.pyccached bytecodeeval loopceval.c, adaptive specialisationreference countingfree when count hits 0GILone thread runs bytecode at a time
swipe the figure sideways, or tap expand for full screen
1/5
compile to bytecode
CPython compiles source to bytecode once and caches it in __pycache__. dis.dis(fn) shows the instructions.
source → bytecode, cached as .pycimport dis; dis.dis(fee)
5

Memory in practice

topicwhat to know
object overheada small int is ~28 bytes; a dict per instance costs memory: use __slots__ or slotted dataclasses for millions of objects
pymallocsmall objects come from arenas; freed memory may not return to the OS, so RSS stays high after a spike
leaksreferences kept in module-level caches, closures, or cycles with __del__; find them with tracemalloc (Backend Debugging part 3)
forked workersGunicorn forks share pages copy-on-write until reference counts change them; gc.freeze() before forking reduces copying
C extensionsNumPy, Pydantic-core (Rust), orjson do heavy work outside the interpreter and can release the GIL