Resources · 101
Progressive releases: stop criteria and rollback
Expand a new version using comparable observations and an exercised recovery path.
· 4 min
What this guide helps achieve
- Define the change and its risks
- Build a useful comparison
- Set expansion conditions
- Exercise complete recovery
- Close the experiment
Quick check
- What starting percentage should be used?
- Does rollback undo a data migration?
- Can expansion proceed without measurements?
Step-by-step method
- 01
Define the change and its risks
Name affected functions, users and dependencies. Separate interface, algorithm, configuration and data changes. List hard-to-reverse effects: notifications, payments, external messages or schema migrations.
Deliverable: scope and irreversible effects.
- 02
Build a useful comparison
Choose a starting population and reference. Check devices, languages, loads and rare tasks. Set an observation period justified by traffic and expected events rather than a universal percentage.
Deliverable: documented population and reference.
- 03
Set expansion conditions
Define errors, latency, completeness and user outcomes with sources and thresholds. Specify who decides to expand, wait or stop. When collection is unavailable, do not treat missing signals as success.
Deliverable: criteria, measurements and decision authority.
- 04
Exercise complete recovery
Test returning to earlier code with current data. Check schema compatibility, queues, caches and ongoing tasks. Earlier software does not undo effects already produced; plan reconciliation and compensating actions where needed.
Deliverable: recovery result and remaining effects.
- 05
Close the experiment
After expansion, check lightly exposed groups and deferred tasks. Remove temporary settings or assign ownership to prevent accumulating variants. Keep versions, decisions and any incident in the change record.
Deliverable: review and temporary-setting cleanup.
Expand, wait or stop
Decisions rely on observations and a prepared recovery path.
Acceptable results
Expand under the plan and inspect lightly exposed groups.
Missing or ambiguous measurements
Wait and restore useful observation before concluding.
Stop criterion reached
Pause expansion, initiate recovery and assess remaining effects.
Illustrative case: rising errors in a lightly represented language may justify a pause even when the overall average appears unchanged.
Choose between continuing, waiting and recovering
Illustrative cases without universal thresholds. Define the criterion before testing, connect it to useful measurements and assign a decision owner.
| Observation | Decision to prepare | Required evidence |
|---|---|---|
| The overall average is stable but one language fails | Pause expansion of the affected segment and inspect its coverage; do not hide failure inside the aggregate. | Results by language, device, task and version, with actually observed counts. |
| The dashboard stops receiving events | Apply the wait condition and restore observation before expanding. | Collection test, freshness and reconciliation of expected versus received operations. |
| The previous version cannot read the new schema | Do not declare rollback ready; test compatibility, restoration or corrected recovery in an isolated environment. | Scenario results with current data, operation order and remaining effects. |
| Notifications have already been sent | Stop the cause, then decide on correction or compensation; reverting code does not withdraw messages. | Proportionate effect list, task state, accountable decision and final reconciliation. |
Reusable worksheet
Complete with your authorised observations. These fields are a working template, not observed results.
| Field | Information to record |
|---|---|
| Exposure | Population, reference, languages and period |
| Decision | Measurements, criteria, owner and stop |
| Recovery | Target version, compatibility, queues and remaining effects |
| Acceptance | Result by scenario, retained evidence and blocking discrepancy |
Worked example
Illustrative situation
Fictional example: the overall metric is stable after limited activation, but a Japanese journey fails. The aggregate dashboard hides this lightly represented segment.
Decision and expected evidence
Apply the planned stop condition to the affected segment, check its collection and replay the task. Before recovery, test the previous version with the current schema and distinguish data restoration from correction of effects already produced.
Distinguish the mechanisms
| Mechanism | Purpose | Check or limitation |
|---|---|---|
| Limited activation | Observe part of the audience | Check representativeness |
| Version rollback | Restore code or configuration | Data may be incompatible |
| Compensating action | Address an effect already produced | Does not guarantee exact restoration |
Management indicators
| Indicator | What it measures | First action |
|---|---|---|
| Observability | Scenarios with a verifiable outcome | Address missing measurements |
| Recovery | Scenarios returned to an acceptable state | Fix dependencies and persistent effects |
Common pitfalls
- Check representativeness
- Data may be incompatible
- Does not guarantee exact restoration
Frequently asked questions
What starting percentage should be used?
There is no universal value. Choose enough exposure to observe useful events while limiting failure consequences.
Does rollback undo a data migration?
Not automatically. Test compatibility and restoration separately, then document remaining effects.
Can expansion proceed without measurements?
Missing measurements do not prove success. Apply the planned wait condition and restore observation before deciding.
Official references
References consulted: . The method and worksheet propose checks to adapt to your context; they do not constitute certification.






