Part 4 · 1 chapters · ~8 min

A/B Testing and Its Traps

Hypothesis tests and p-values, statistical power and sample size, the peeking problem simulated, multiple comparisons, choosing one primary metric in advance, novelty and seasonality effects, sample ratio mismatch, guardrail metrics, sequential testing, and CUPED variance reduction.

5

Simulated: what peeking does

code
// A/A: both arms convert at 5%; check after every 500 users per arm, up to 10,000
for (let step = 0; step < 20; step++) {
  for (let i = 0; i < 500; i++) { ca += rnd() < 0.05; cb += rnd() < 0.05; } n += 500;
  const p = (ca + cb) / (2 * n), se = Math.sqrt(2 * p * (1 - p) / n), z = (cb / n - ca / n) / se;
  if (Math.abs(z) > 1.96) { stopped = true; break; }          // "significant!": ship it
}
// 2,000 experiments: peeking 24.1% false positives · fixed horizon 4.7%

// sample size per arm to detect 5.0% → 5.5% (α = 0.05 two-sided, power 0.8): 31,234
trapdefence
peeking and stopping earlyfixed sample size, or a sequential method
testing 20 metrics and reporting the winnerone primary metric chosen before launch; correct for multiple comparisons
sample ratio mismatch (50/50 split gives 52/48)a chi-square check; a mismatch means the assignment or logging is broken
novelty effectsrun at least one or two full weekly cycles
winning on clicks, losing on revenueguardrail metrics that must not get worse
A/A TESTS: NO REAL DIFFERENCE
2,000 simulated experiments, 5% conversion in both arms, 10,000 users per arm
check once at the planned end4.7% false positivespeek 20 times, stop at first p < 0.0524.1% false positives
swipe the figure sideways, or tap expand for full screen
1/4
the setup
Both arms are identical, so every "significant" result is a false positive. A correctly run test should produce about 5%.
no true effectexpect ~5%