Part 5 · 2 chapters · ~20 min
Experimentation and Flags
The experiment platform as four services (deterministic assignment with layers, exposure logging as the denominator, company-wide metrics, analysis with variance reduction and SRM checks) and the client's four jobs, then holdouts that measure what every launch did together, guardrails as a floor, and the launch review where people who did not run the experiment decide.
11
The experiment platform
four services
- Assignment: variant from a hash of the unit id and the experiment's salt, mapped by allocation: deterministic (the same user on every device without a lookup), sticky, independent per experiment. Evaluated on the client from a bootstrapped config, or once per session from a service.
- Layers: experiments on the same surface are mutually exclusive within a layer (a unit is in at most one per layer); layers run independently on the same users; capacity within a layer is a scheduling problem the platform owns. The collision answer.
- Exposure: assignment is not exposure; a user assigned to a new checkout who never visits it has no information. The client logs exposure the first time the unit encounters the variant, reliably (a beacon that survives slow devices), and analysis uses exposed units as the population. Logging exposure at assignment, or failing to log on slow clients, are the two most common invalid experiments and the source of sample-ratio mismatch.
- Metrics: defined once in a company metrics repository (conversion, revenue per user, p75 LCP and INP, crash-free sessions, day-7 retention), computed per unit and joined to exposures. A team pre-registers a primary and declares guardrails; nobody defines their own "conversion".
- Analysis: confidence intervals; variance reduction (CUPED from pre-experiment behaviour) to detect small effects with fewer units; sequential tests or fixed horizons against peeking; the SRM check first; multiple-comparison correction. A decision-ready report, not a p-value.
- The client's four jobs: evaluate locally without a round trip (no flicker); log exposure at encounter with a beacon; tag every event with experiment and variant; never leak a variant into cached or shared state (a server-rendered page cached with one user's variant is served to everyone).
THE EXPERIMENT PLATFORM
assignment, exposure, metrics, analysis: the four services, and what each one must get right
swipe the figure sideways, or tap expand for full screen
1/6
assignment
Assignment: variant = hash(unitId + experiment.salt) mod buckets, mapped to variants by allocation. Deterministic, so the same user sees the same variant on every device and every day without a lookup; sticky, so a user is never switched mid-experiment; independent per experiment because the salt differs, so being in A for one experiment says nothing about another. The client evaluates it locally from a bootstrapped config, or asks a service once per session.
12
Holdouts, guardrails and the launch review
what a year of launches did together
- The holdout: 1 to 5% of users kept on the old experience for a period, excluded from every launch by the platform, compared against the rest at the end. A quarter of launches that summed to +12% individually measured +4% in the holdout; the rest was cannibalisation and novelty. The only instrument that measures interactions between launches.
- Mechanics: the same deterministic hash with a holdout salt; launched flags are "on for everyone except H"; refreshed each period; the holdout is a real product for real users (security fixes and infrastructure changes apply).
- Guardrails: a list per surface owned above the team (crash-free sessions, error rate, p75 LCP and INP by device tier, revenue per user, support contacts, accessibility pass rate), each a non-inferiority test with a tolerance. A primary win that trips a guardrail is a redesign, not a launch.
- The launch review: the pre-registered primary with its interval, every guardrail with its interval, the SRM check, exposure counts, duration against the pre-computed minimum, and segments (country, device tier, new versus returning: a desktop win that loses on 2 GB phones is not a win). Reviewers who did not run the experiment decide ship, iterate or stop, and the decision is recorded with reasons.
- The failure modes it catches: a metric chosen after the result; a test stopped on its first significant day; a segment win sold as global; a guardrail explained away; an atypical week without a rerun. Adversarial by design, polite by culture.
- Costs and buys: 2% of users without improvements for a quarter (argued about every quarter), slower launches, a meeting with preparation; against knowing what the launches did in total, a floor that cannot be lowered one experiment at a time, and no manufactured wins. At a thousand experiments a year, the difference between a product that improves and one that churns.
the transferable part
Pre-register the metric, compute the minimum duration before starting, check the sample ratio first, declare guardrails, and have someone else decide. A team of thirty with a flag service and a spreadsheet can do all five; the platform is what makes them automatic at a thousand.
HOLDOUTS, GUARDRAILS AND THE LAUNCH REVIEW
measuring what a year of launches did together, the floor nobody may lower, and the meeting where numbers decide
swipe the figure sideways, or tap expand for full screen
1/6
the holdout
The holdout: at the start of a period, 2% of users are assigned to a holdout that receives no launches from a set of teams (the feed team's holdout; the growth org's holdout); they see the product as it was. At the end of the period, the holdout versus the rest measures the sum of every launch that period, including interactions the individual experiments could not see. A quarter of launches that individually summed to +12% engagement measured +4% in the holdout: the rest was cannibalisation and novelty.