Part 5 · 1 chapters · ~8 min
Inference Serving: KV Cache, Batching, Quantisation
Prefill versus decode, why decoding is memory-bandwidth bound, the KV cache and its arithmetic, continuous batching, PagedAttention, prefix caching, speculative decoding, quantisation (int8, 4-bit, GGUF, AWQ) and its quality cost, serving stacks (vLLM, TGI, SGLang, llama.cpp), and the metrics that matter (time to first token, tokens per second, cost per million tokens).
6
Prefill, decode and memory
code
prefill: process the whole prompt in parallel → compute-bound, sets time to first token (TTFT)
decode: one token at a time, each step reads every weight once → memory-bandwidth bound
weights: 8B params × 2 bytes (fp16) = 16.1 GB · int8 8.0 GB · 4-bit 4.0 GB (+ overhead)
70B: 141 GB fp16 · 35 GB 4-bit
decode speed bound ≈ memory bandwidth / bytes read per token
Apple M3 Pro, 150 GB/s, 8B at ~4.5 GB in 4-bit → at most ~33 tokens/s for one sequence
KV cache per token = 2 × 32 layers × 8 KV heads × 128 × 2 bytes = 128 KiB (Llama 3.1 8B)| technique | effect |
|---|---|
| continuous batching | add and remove sequences each step: many users share each weight read, raising throughput |
| PagedAttention (vLLM) | allocate KV cache in pages: little waste, more sequences per GPU |
| prefix caching | reuse the KV cache of a shared system prompt across requests (provider prompt caching works this way) |
| speculative decoding | a small draft model proposes tokens, the large model verifies several at once |
| quantisation | fewer bytes per weight: faster decode, more room; small quality loss at 8 bits, more at 4 |
KV CACHE MEMORY, COMPUTED
2 × layers × KV heads × head dim × 2 bytes (fp16) per token
swipe the figure sideways, or tap expand for full screen
1/4
why a cache
Generating each new token needs the keys and values of every earlier token. Recomputing them each step would be quadratic, so servers cache them: the KV cache.
keep K and V of past tokensavoid recomputation