Part 8 · 1 chapters · ~8 min

Performance and Profiling

Profiling with cProfile, py-spy and Scalene, the N+1 hunt with django-debug-toolbar and nplusone, database indexes and query plans from Django, caching with Django's cache framework and Redis, serialisation costs (orjson), vectorising with NumPy, and when to reach for Rust or C extensions.

14

Where Python services spend time

code
py-spy record -o flame.svg --pid $(pgrep -f gunicorn | head -1) --duration 30   # production-safe, no restart
python -m cProfile -o out.prof manage.py reconcile 2026-09-30 && snakeviz out.prof
scalene script.py                    # CPU, memory and GPU per line; separates Python time from native time

# caching with Django's framework, Redis backend
from django.core.cache import cache
rate = cache.get_or_set(f"rate:{pair}", lambda: fetch_rate(pair), timeout=30)
usual bottleneck in Django servicesfix
N+1 queriesselect_related / prefetch_related; nplusone in tests
missing indexesMeta.indexes, check plans with qs.explain(analyze=True)
serialising large responsespaginate; orjson; DRF serializers are slow for big lists, use .values() for read-only endpoints
CPU-heavy loopsNumPy or Polars vectorisation; move to Celery workers; Rust extensions via PyO3 for hot kernels
synchronous external calls in requestsCelery, or async views with httpx