Part 4 · 1 chapters · ~8 min
A/B Testing and Its Traps
Hypothesis tests and p-values, statistical power and sample size, the peeking problem simulated, multiple comparisons, choosing one primary metric in advance, novelty and seasonality effects, sample ratio mismatch, guardrail metrics, sequential testing, and CUPED variance reduction.
5
Simulated: what peeking does
code
// A/A: both arms convert at 5%; check after every 500 users per arm, up to 10,000
for (let step = 0; step < 20; step++) {
for (let i = 0; i < 500; i++) { ca += rnd() < 0.05; cb += rnd() < 0.05; } n += 500;
const p = (ca + cb) / (2 * n), se = Math.sqrt(2 * p * (1 - p) / n), z = (cb / n - ca / n) / se;
if (Math.abs(z) > 1.96) { stopped = true; break; } // "significant!": ship it
}
// 2,000 experiments: peeking 24.1% false positives · fixed horizon 4.7%
// sample size per arm to detect 5.0% → 5.5% (α = 0.05 two-sided, power 0.8): 31,234| trap | defence |
|---|---|
| peeking and stopping early | fixed sample size, or a sequential method |
| testing 20 metrics and reporting the winner | one primary metric chosen before launch; correct for multiple comparisons |
| sample ratio mismatch (50/50 split gives 52/48) | a chi-square check; a mismatch means the assignment or logging is broken |
| novelty effects | run at least one or two full weekly cycles |
| winning on clicks, losing on revenue | guardrail metrics that must not get worse |
A/A TESTS: NO REAL DIFFERENCE
2,000 simulated experiments, 5% conversion in both arms, 10,000 users per arm
swipe the figure sideways, or tap expand for full screen
1/4
the setup
Both arms are identical, so every "significant" result is a false positive. A correctly run test should produce about 5%.
no true effectexpect ~5%