7 parts · 7 chapters

Linear Algebra and the Maths of ML

The maths under embeddings, vector search and neural networks, taught with code that runs. In the course's gradient-descent example, a learning rate of 0.01 recovered the true line (w = 3.001, b = 1.990), 0.001 was still converging after 2,000 steps, and 0.03 diverged to 10⁸¹.

Seven parts: vectors and matrices; similarity and embeddings (cosine similarity, vector search, and why a million 1,536-dimension embeddings take 6.14 GB); linear maps; eigenvectors and PCA; calculus for gradients; optimisation and gradient descent; and probability for models (softmax, cross-entropy, likelihood).

vectors and matrices · similarity and embeddings · linear maps · eigenvectors · calculus for gradients · optimisation and gradient descent · probability for modelsbeginner → senior · engineers building with AI
vectorsLists of numbers with length and direction; dot products.
embeddingsMeaning as position; similarity as angle.
matricesFunctions that transform vectors: layers of a network.
eigenDirections a matrix only stretches: PCA, PageRank.
gradientsWhich way to nudge every parameter to reduce loss.
probabilitySoftmax, cross-entropy and likelihood: how models are scored.
For programmers without a CS degreeUses Probability and Statistics; leads into How LLMs Work, AI-native engineering and the Search course's vector search part.