Part 6 · 1 chapters · ~8 min

Probability for Models

Models that output probabilities, logits and softmax (computed), numerical stability, cross-entropy and log-likelihood, maximum likelihood training, temperature and sampling, calibration, classification thresholds and precision versus recall, and evaluating models with held-out data.

7

From scores to decisions

code
const softmax = (z: number[]) => { const m = Math.max(...z); const e = z.map(t => Math.exp(t - m)); const s = e.reduce((a, c) => a + c); return e.map(t => t / s); };
softmax([2, 1, 0.1])            // [0.659, 0.242, 0.099]  (subtracting the max avoids overflow)
-Math.log(softmax([2, 1, 0.1])[0])   // 0.417: cross-entropy if class 0 is correct
softmax([2, 1, 0.1].map(z => z / 0.5))   // temperature 0.5: sharper, more confident

// a language model: softmax over ~100k tokens at every step; training minimises the average
// -log p(actual next token) over trillions of tokens (How LLMs Work course)

Calibration: a model is calibrated if, among predictions made with 80% confidence, about 80% are right. Thresholds for actions (block, review, allow) should come from precision and recall at the base rates you face (Statistics P7, the fraud example), not from the raw score.

SOFTMAX AND CROSS-ENTROPY, COMPUTED
scores [2, 1, 0.1] turned into probabilities
class 0 (score 2)0.659class 1 (score 1)0.242class 2 (score 0.1)0.099
swipe the figure sideways, or tap expand for full screen
1/4
scores to probabilities
Softmax exponentiates scores and divides by the sum, giving positive numbers that sum to 1: [2, 1, 0.1] becomes [0.659, 0.242, 0.099].
exp, then normalisesums to 1