Resources · 26

Run an A/B experiment that supports a sound decision

Specify hypothesis, assignment, primary metric and guardrails before comparing variants.

· 19 min

Product team reviewing an A/B experiment

What this guide helps achieve

  • Preregister the protocol
  • Check assignment
  • Watch guardrails
  • Decide with uncertainty

Quick check

  • What precise hypothesis could be disproved?
  • What unit receives the variant?
  • Does the observed allocation match the plan?
  • Which harm warrants a pause?
  • Were duration and segments set before results were seen?

Step-by-step method

  1. 01

    State the proposed change

    Write the expected mechanism, target population and effect on one primary metric. Describe possible accessibility and service risks.

    Deliverable: dated hypothesis and success criteria.

  2. 02

    Define assignment

    Choose the randomization unit and keep assignment stable across visits. Exclude bots and specify out-of-scope segments before launch.

    Deliverable: assignment protocol and exclusions.

  3. 03

    Instrument and test

    Verify exposure, conversion and guardrail events for both variants. Test mobile, consent refusal and error states.

    Deliverable: QA record of both journeys.

  4. 04

    Check sample ratio

    Compare observed and expected allocations before reading the treatment effect. A mismatch may expose assignment, measurement or targeting defects.

    Deliverable: ratio check and discrepancy diagnosis.

  5. 05

    Analyze at the planned time

    Respect the chosen window and report uncertainty, primary metric, harms and planned segment analysis. Inspect surprises without selecting only the favorable number after the fact.

    Deliverable: analysis with limitations and uncertainty.

  6. 06

    Decide and follow up

    Assign an owner to rollout, abandonment or another test. After rollout, check that effect and guardrails remain consistent.

    Deliverable: archived decision and post-rollout check.

Management indicators

IndicatorWhat it measuresFirst action
AssignmentExpected versus observed sample ratioDiagnose before analysis
ExposureEligible users actually exposedRepair targeting
EffectPrimary-metric change and uncertaintyApply planned decision rule
GuardrailsErrors, latency and abandonment per variantPause harmful treatment

Common pitfalls

  • Stopping when a number first looks favorable
  • Collecting many metrics without a primary one
  • Ignoring sample ratio mismatch
  • Treating a tiny segment as general proof

Frequently asked questions

Is sample ratio a minor formality?

No. An unexplained mismatch can make the comparison unreliable.

Should conversion be checked daily?

Operational monitoring helps, but changing duration or the decision rule in response to results undermines the planned interpretation.

What if the result is uncertain?

Report uncertainty and plausible effect size, then decide whether another experiment has real value.

Official references