Part 2 · 1 chapters · ~8 min
Linear Maps
Linear functions and matrices as the same thing, columns as images of basis vectors, scaling, rotation, shear and projection, composition by multiplication, inverses and determinants, rank, why neural networks need non-linearities, and attention as matrix products.
3
Transformations in code
code
const rot90 = [[Math.cos(Math.PI / 2), -Math.sin(Math.PI / 2)], [Math.sin(Math.PI / 2), Math.cos(Math.PI / 2)]]; matvec(rot90, [1, 0]) // [0, 1] (after rounding floating-point noise) // a two-layer network: linear, non-linear, linear const relu = (v: number[]) => v.map(x => Math.max(0, x)); const layer = (W: number[][], b: number[], x: number[]) => matvec(W, x).map((y, i) => y + b[i]); const forward = (x: number[]) => layer(W2, b2, relu(layer(W1, b1, x))); // without relu: W2(W1 x + b1) + b2 = (W2 W1) x + (W2 b1 + b2): one linear map, no extra power # attention (How LLMs Work course), as matrix products: softmax(Q Kᵀ / √d) V
Determinant and inverse: the determinant says how a matrix scales area or volume; zero means it flattens space and has no inverse (information is lost). Rank counts the independent directions a matrix can output; low-rank approximations underlie LoRA fine-tuning, which trains small rank-r updates instead of whole matrices.
MATRICES ARE FUNCTIONS
what a 2 × 2 matrix does to the plane
swipe the figure sideways, or tap expand for full screen
1/4
linear
A linear map preserves addition and scaling: f(a + b) = f(a) + f(b). Every such map on vectors is a matrix, and every matrix is such a map.
matrices = linear functionsf(a+b) = f(a)+f(b)