Resources · 26
Run an A/B experiment that supports a sound decision
Specify hypothesis, assignment, primary metric and guardrails before comparing variants.
· 19 min
What this guide helps achieve
- Preregister the protocol
- Check assignment
- Watch guardrails
- Decide with uncertainty
Quick check
- What precise hypothesis could be disproved?
- What unit receives the variant?
- Does the observed allocation match the plan?
- Which harm warrants a pause?
- Were duration and segments set before results were seen?
Step-by-step method
- 01
State the proposed change
Write the expected mechanism, target population and effect on one primary metric. Describe possible accessibility and service risks.
Deliverable: dated hypothesis and success criteria.
- 02
Define assignment
Choose the randomization unit and keep assignment stable across visits. Exclude bots and specify out-of-scope segments before launch.
Deliverable: assignment protocol and exclusions.
- 03
Instrument and test
Verify exposure, conversion and guardrail events for both variants. Test mobile, consent refusal and error states.
Deliverable: QA record of both journeys.
- 04
Check sample ratio
Compare observed and expected allocations before reading the treatment effect. A mismatch may expose assignment, measurement or targeting defects.
Deliverable: ratio check and discrepancy diagnosis.
- 05
Analyze at the planned time
Respect the chosen window and report uncertainty, primary metric, harms and planned segment analysis. Inspect surprises without selecting only the favorable number after the fact.
Deliverable: analysis with limitations and uncertainty.
- 06
Decide and follow up
Assign an owner to rollout, abandonment or another test. After rollout, check that effect and guardrails remain consistent.
Deliverable: archived decision and post-rollout check.
Management indicators
| Indicator | What it measures | First action |
|---|---|---|
| Assignment | Expected versus observed sample ratio | Diagnose before analysis |
| Exposure | Eligible users actually exposed | Repair targeting |
| Effect | Primary-metric change and uncertainty | Apply planned decision rule |
| Guardrails | Errors, latency and abandonment per variant | Pause harmful treatment |
Common pitfalls
- Stopping when a number first looks favorable
- Collecting many metrics without a primary one
- Ignoring sample ratio mismatch
- Treating a tiny segment as general proof
Frequently asked questions
Is sample ratio a minor formality?
No. An unexplained mismatch can make the comparison unreliable.
Should conversion be checked daily?
Operational monitoring helps, but changing duration or the decision rule in response to results undermines the planned interpretation.
What if the result is uncertain?
Report uncertainty and plausible effect size, then decide whether another experiment has real value.






