Part 5 · 1 chapters · ~8 min

Inference Serving: KV Cache, Batching, Quantisation

Prefill versus decode, why decoding is memory-bandwidth bound, the KV cache and its arithmetic, continuous batching, PagedAttention, prefix caching, speculative decoding, quantisation (int8, 4-bit, GGUF, AWQ) and its quality cost, serving stacks (vLLM, TGI, SGLang, llama.cpp), and the metrics that matter (time to first token, tokens per second, cost per million tokens).

6

Prefill, decode and memory

code
prefill: process the whole prompt in parallel → compute-bound, sets time to first token (TTFT)
decode:  one token at a time, each step reads every weight once → memory-bandwidth bound

weights: 8B params × 2 bytes (fp16) = 16.1 GB · int8 8.0 GB · 4-bit 4.0 GB (+ overhead)
         70B: 141 GB fp16 · 35 GB 4-bit
decode speed bound ≈ memory bandwidth / bytes read per token
         Apple M3 Pro, 150 GB/s, 8B at ~4.5 GB in 4-bit → at most ~33 tokens/s for one sequence

KV cache per token = 2 × 32 layers × 8 KV heads × 128 × 2 bytes = 128 KiB (Llama 3.1 8B)
techniqueeffect
continuous batchingadd and remove sequences each step: many users share each weight read, raising throughput
PagedAttention (vLLM)allocate KV cache in pages: little waste, more sequences per GPU
prefix cachingreuse the KV cache of a shared system prompt across requests (provider prompt caching works this way)
speculative decodinga small draft model proposes tokens, the large model verifies several at once
quantisationfewer bytes per weight: faster decode, more room; small quality loss at 8 bits, more at 4
KV CACHE MEMORY, COMPUTED
2 × layers × KV heads × head dim × 2 bytes (fp16) per token
Llama 3.1 8B, per token128 KiB8B at 8k context1.00 GiB70B at 8k context2.50 GiB8B at 128k context16.0 GiB
swipe the figure sideways, or tap expand for full screen
1/4
why a cache
Generating each new token needs the keys and values of every earlier token. Recomputing them each step would be quadratic, so servers cache them: the KV cache.
keep K and V of past tokensavoid recomputation