8 parts · 8 chapters
How LLMs Work
What happens between your prompt and the reply, with the arithmetic that decides cost and speed. Measured with OpenAI's o200k_base tokenizer: the Yoruba greeting "Ẹ káàárọ̀, ṣé dáadáa ni?" is 15 tokens for 24 characters, while "Good morning, how are you?" is 7 tokens for 26. Computed: an 8B-parameter model's KV cache costs 128 KiB per token, 16 GiB at a 128k context.
Eight parts: tokenisation; embeddings; attention and transformers; training (pre-training, instruction tuning, preference optimisation); fine-tuning with LoRA; inference serving (prefill and decode, the KV cache, batching, quantisation); evaluation; and running a model locally, with the memory-bandwidth arithmetic that predicts its speed.
tokensText becomes integer ids; languages and digits tokenise unevenly.
vectorsTokens become embeddings; meaning becomes geometry.
attentionEach token mixes in information from the tokens before it.
trainingNext-token prediction, then instruction and preference tuning.
servingMemory, not compute, limits decoding speed.
evaluationYour own task-specific evals beat leaderboards.
00
Tokenisation
Text to integers
1 ch · ~8 min01Embeddings
A lookup, then geometry
1 ch · ~8 min02Attention and Transformers
Attention in a few lines
1 ch · ~8 min03Training
The arithmetic of training
1 ch · ~8 min04Fine-Tuning with LoRA
Adapters in code
1 ch · ~8 min05Inference Serving: KV Cache, Batching, Quantisation
Prefill, decode and memory
1 ch · ~8 min06Evaluation
An eval harness
1 ch · ~8 min07Running a Model Locally
Commands and arithmetic
1 ch · ~8 minBuilt on the ML maths courseUses Linear Algebra and the Maths of ML for vectors, matrices, softmax and gradients; feeds the AI-native engineering course.