Part 3 · 1 chapters · ~8 min
Training
Pre-training objective and data, scaling laws (Kaplan 2020, Chinchilla 2022: about 20 tokens per parameter as a compute-optimal reference), the compute budget in FLOPs (about 6 × parameters × tokens), supervised fine-tuning, RLHF and DPO, reinforcement learning with verifiable rewards, data quality and contamination, and why training is out of reach for most teams while fine-tuning is not.
4
The arithmetic of training
code
training compute ≈ 6 × N (parameters) × D (tokens) FLOPs 8B parameters × 15T tokens (Meta reported ~15T tokens for Llama 3) ≈ 6 × 8e9 × 15e12 = 7.2e23 FLOPs at 400 TFLOP/s effective per GPU: 7.2e23 / 4e14 ≈ 1.8e9 GPU-seconds ≈ 500,000 GPU-hours Chinchilla (Hoffmann et al. 2022): for a fixed compute budget, scale parameters and tokens together, roughly 20 tokens per parameter. Modern small models train far beyond that (Llama 3 8B: ~1,900 tokens/param) because a smaller model trained longer is cheaper to serve.
Data decides behaviour: deduplication, filtering, the mix of code, maths and languages, and benchmark contamination (test questions leaking into training data) matter as much as architecture. The amount of Yoruba text in pre-training is one reason the tokenizer and model treat Yoruba less efficiently (part 1).
FROM RANDOM WEIGHTS TO AN ASSISTANT
the stages of training a chat model
swipe the figure sideways, or tap expand for full screen
1/4
pre-training
The model learns to predict the next token over trillions of tokens of text and code by minimising cross-entropy (ML maths P7). Almost all compute is spent here; the result completes text but does not follow instructions well.
next-token prediction at scalemost of the compute