Part 4 · 1 chapters · ~8 min
Calculus for Gradients
Derivatives as rates of change, numerical derivatives and checking them, rules for common functions, partial derivatives, the gradient vector, the chain rule, computational graphs, backpropagation and automatic differentiation (PyTorch autograd), and vanishing and exploding gradients.
5
Gradients by hand and by machine
code
// numeric derivative: a test for any analytic gradient const f = (t: number) => t * t, h = 1e-5; (f(3 + h) - f(3 - h)) / (2 * h) // 6.000000 = 2·3 // loss for one point and its gradient by the chain rule // L = (w x + b - y)² → ∂L/∂w = 2 (w x + b - y) x , ∂L/∂b = 2 (w x + b - y) # PyTorch computes the same automatically import torch w = torch.tensor(0.0, requires_grad=True); b = torch.tensor(0.0, requires_grad=True) loss = ((w * x + b - y) ** 2).mean() loss.backward() # fills w.grad and b.grad via backpropagation print(w.grad, b.grad)
Vanishing and exploding gradients: multiplying many local derivatives through deep networks can shrink gradients to nothing or blow them up. Residual connections, normalisation layers and careful initialisation exist to keep them in a usable range.
DERIVATIVES AND THE CHAIN RULE
how a change in each parameter changes the loss
swipe the figure sideways, or tap expand for full screen
1/4
derivative
A derivative is the rate of change: how much the output moves per tiny change in the input. Numerically, (f(3 + h) - f(3 - h)) / 2h for f(x) = x² gives 6.000000, matching the rule d/dx x² = 2x.
rate of changenumeric check: 6.000000