Part 8 · 1 chapters · ~8 min
Performance and Profiling
Profiling with cProfile, py-spy and Scalene, the N+1 hunt with django-debug-toolbar and nplusone, database indexes and query plans from Django, caching with Django's cache framework and Redis, serialisation costs (orjson), vectorising with NumPy, and when to reach for Rust or C extensions.
14
Where Python services spend time
code
py-spy record -o flame.svg --pid $(pgrep -f gunicorn | head -1) --duration 30 # production-safe, no restart
python -m cProfile -o out.prof manage.py reconcile 2026-09-30 && snakeviz out.prof
scalene script.py # CPU, memory and GPU per line; separates Python time from native time
# caching with Django's framework, Redis backend
from django.core.cache import cache
rate = cache.get_or_set(f"rate:{pair}", lambda: fetch_rate(pair), timeout=30)| usual bottleneck in Django services | fix |
|---|---|
| N+1 queries | select_related / prefetch_related; nplusone in tests |
| missing indexes | Meta.indexes, check plans with qs.explain(analyze=True) |
| serialising large responses | paginate; orjson; DRF serializers are slow for big lists, use .values() for read-only endpoints |
| CPU-heavy loops | NumPy or Polars vectorisation; move to Celery workers; Rust extensions via PyO3 for hot kernels |
| synchronous external calls in requests | Celery, or async views with httpx |