8 parts · 8 chapters

How LLMs Work

What happens between your prompt and the reply, with the arithmetic that decides cost and speed. Measured with OpenAI's o200k_base tokenizer: the Yoruba greeting "Ẹ káàárọ̀, ṣé dáadáa ni?" is 15 tokens for 24 characters, while "Good morning, how are you?" is 7 tokens for 26. Computed: an 8B-parameter model's KV cache costs 128 KiB per token, 16 GiB at a 128k context.

Eight parts: tokenisation; embeddings; attention and transformers; training (pre-training, instruction tuning, preference optimisation); fine-tuning with LoRA; inference serving (prefill and decode, the KV cache, batching, quantisation); evaluation; and running a model locally, with the memory-bandwidth arithmetic that predicts its speed.

tokenisation · embeddings · attention and transformers · training · fine-tuning (LoRA) · inference serving (KV cache, batching, quantisation) · evaluation · running a model locallymid → staff · engineers building with or on LLMs
tokensText becomes integer ids; languages and digits tokenise unevenly.
vectorsTokens become embeddings; meaning becomes geometry.
attentionEach token mixes in information from the tokens before it.
trainingNext-token prediction, then instruction and preference tuning.
servingMemory, not compute, limits decoding speed.
evaluationYour own task-specific evals beat leaderboards.
Built on the ML maths courseUses Linear Algebra and the Maths of ML for vectors, matrices, softmax and gradients; feeds the AI-native engineering course.