Part 5 · 1 chapters · ~8 min

Optimisation and Gradient Descent

Loss functions, gradient descent, learning rates that are too small, right and too large (computed), stochastic and mini-batch gradient descent, momentum and Adam, learning-rate schedules and warm-up, local minima and saddle points, overfitting, regularisation and validation sets.

6

Three learning rates, computed

code
let w = 0, b = 0;
for (let it = 0; it <= 2000; it++) {
  let gw = 0, gb = 0;
  for (let i = 0; i < xs.length; i++) { const e = w * xs[i] + b - ys[i]; gw += 2 * e * xs[i]; gb += 2 * e; }
  w -= lr * gw / xs.length; b -= lr * gb / xs.length;
}
// data: y = 3x + 2 + noise, 100 points, x in [0, 10)
// lr 0.001 → w 3.094 b 1.385  (loss 364.7 → 0.568 at step 100 → 0.180 at 2,000)
// lr 0.01  → w 3.001 b 1.990  (loss 0.304 at 100 → 0.095 at 1,000)
// lr 0.03  → diverges: 899.6 at step 10, 3,071,544 at 100, ~1.2e81 at 2,000
techniquewhat it fixes
mini-batch SGDfull-dataset gradients are too expensive; batches give noisy but cheap estimates
momentum, Adamsmooth noisy steps and adapt the step per parameter; Adam is the default for transformers
warm-up and decay scheduleslarge early steps diverge; small late steps refine
validation set, early stopping, weight decayoverfitting: low training loss, poor results on new data
GRADIENT DESCENT: THREE LEARNING RATES
fitting y = 3x + 2 (+ noise) to 100 points; mean squared error after 100 steps
lr 0.0010.568 (slow)lr 0.010.304, then 0.095 at step 1,000lr 0.033,071,544 (diverging)
swipe the figure sideways, or tap expand for full screen
1/4
the loop
Gradient descent repeats: compute the gradient of the loss, step each parameter a little against it. The step size is the learning rate.
step against the gradientw ← w - lr × ∂L/∂w