Part 6 · 2 chapters · ~12 min
Experiments and Launches
Hypotheses with predicted effects, A/B tests and their pitfalls (peeking, novelty effects, sample ratio mismatch), when not to experiment, staged launches from dogfood to GA, launch checklists across support, legal, marketing and operations, and post-launch reviews.
12
From hypothesis to launch
code
hypothesis: if borrowers can generate their clearance letter in the app,
then manual letter tickets will fall by at least 50% within 6 weeks of GA,
because 80% of tickets are simple requests from eligible borrowers.
we will know we are wrong if tickets fall less than 20%.A STAGED LAUNCH
from internal dogfood to general availability
swipe the figure sideways, or tap expand for full screen
1/5
dogfood
Staff use it first: obvious bugs and confusing copy surface before customers see them.
staff firstcheap bugs found cheaply
13
Experiment pitfalls and launch checklists
| pitfall | effect | guard |
|---|---|---|
| peeking and stopping when significant | false positives far above the stated rate | fix sample size and duration upfront, or use sequential methods |
| novelty effect | a short-term bump that fades | run long enough; look at returning users |
| sample ratio mismatch | broken randomisation invalidates results | check the split matches the design before reading results |
| too many metrics | something will look significant by chance | one primary metric, declared in advance |
When not to experiment: regulatory requirements, obvious bug fixes, tiny audiences that can never reach significance, and changes where withholding the improvement from a control group is unfair (for example, a fraud protection).