7 parts · 7 chapters
Linear Algebra and the Maths of ML
The maths under embeddings, vector search and neural networks, taught with code that runs. In the course's gradient-descent example, a learning rate of 0.01 recovered the true line (w = 3.001, b = 1.990), 0.001 was still converging after 2,000 steps, and 0.03 diverged to 10⁸¹.
Seven parts: vectors and matrices; similarity and embeddings (cosine similarity, vector search, and why a million 1,536-dimension embeddings take 6.14 GB); linear maps; eigenvectors and PCA; calculus for gradients; optimisation and gradient descent; and probability for models (softmax, cross-entropy, likelihood).
vectorsLists of numbers with length and direction; dot products.
embeddingsMeaning as position; similarity as angle.
matricesFunctions that transform vectors: layers of a network.
eigenDirections a matrix only stretches: PCA, PageRank.
gradientsWhich way to nudge every parameter to reduce loss.
probabilitySoftmax, cross-entropy and likelihood: how models are scored.
00
Vectors and Matrices
The operations in code
1 ch · ~8 min01Similarity and Embeddings
Similarity is an angle
1 ch · ~8 min02Linear Maps
Transformations in code
1 ch · ~8 min03Eigenvectors
Power iteration in code
1 ch · ~8 min04Calculus for Gradients
Gradients by hand and by machine
1 ch · ~8 min05Optimisation and Gradient Descent
Three learning rates, computed
1 ch · ~8 min06Probability for Models
From scores to decisions
1 ch · ~8 minFor programmers without a CS degreeUses Probability and Statistics; leads into How LLMs Work, AI-native engineering and the Search course's vector search part.