Resources · 49
Evaluate AI regressions before changing versions
Compare versions on the same tasks, find local losses and make a decision supported by cases rather than an average.
· 3 min
What this guide helps achieve
- Isolate the tested change
- Preserve an independent reference set
- Inspect failures by segment
- Prepare shutdown or rollback
Quick check
- Which component changed?
- Was the test set used to optimise the candidate?
- Who resolves annotation disagreements?
- Does an average hide a sensitive error?
- Can the previous configuration be restored?
Step-by-step method
- 01
Define task and refusals
Describe expected output, permitted data and costly errors. Separate accuracy, usefulness, appropriate refusal and action effects. NIST’s framework calls for context-appropriate measurement; it does not provide a universal reliability score.
Deliverable: criteria and decision owner.
- 02
Separate reference and discovery
Keep a stable set out of optimisation. Add new anonymised feedback cases separately: they reveal blind spots but should not silently change the comparison denominator.
Deliverable: versioned sets with provenance.
- 03
Document annotation
Explain what makes an answer acceptable, partial or incorrect. Have a competent reviewer resolve ambiguity and retain reasons for disagreements. A single reference answer may be too restrictive for an open-ended task.
Deliverable: rubric and adjudicated cases.
- 04
Compare under equal conditions
Replay identical inputs with recorded versions, parameters, tools and sources. Repeat trials when outputs vary and report variation. Separate model failure, absent data, unavailable tools and a defective test.
Deliverable: paired outcomes and minimised traces.
- 05
Decide by risk family
Inspect losses by language, document type and sensitive situation even when the mean improves. Define stop, correction and limited-release criteria before reading results. One severe case can justify withholding a release.
Deliverable: decision, reservations and exceptions.
- 06
Replay after release
Prepare a restorable configuration and reachable owner. Compare early feedback with the test set without treating local success as evidence for all uses. Reassess when data or dependencies change.
Deliverable: monitoring and rollback procedure.
Reusable worksheet
Complete with your authorised observations. These fields are a working template, not observed results.
| Field | Information to record |
|---|---|
| Case and segment | Stable identifier, language, task type |
| Expected | Acceptance criterion and reference evidence |
| Versions A / B | Outcome, repetition, failure cause |
| Decision | Accept, correct or stop; owner and evidence |
Worked example
Illustrative situation
Fictional example: a release improves routine answers but drops uncertainty statements when documents conflict.
Decision and expected evidence
The report isolates that subset, retains source evidence and withholds release for that use until a correction is retested.
Distinguish the mechanisms
| Mechanism | Purpose | Check or limitation |
|---|---|---|
| Stable set | Compare versions | Keep out of optimisation |
| New cases | Find blind spots | Report separately from historical results |
| In-service observation | Understand actual use | Respect permissions, minimisation and context |
Management indicators
| Indicator | What it measures | First action |
|---|---|---|
| Paired losses | Previously accepted cases now failing | Inspect each sensitive loss |
| Annotation disagreement | Cases with differing judgements | Clarify the rubric before comparison |
| Segment coverage | Situations actually represented | Name untested uses |
Common pitfalls
- Optimising on the final test
- Changing set and version together
- Treating an automated judge as truth
- Accepting critical losses because the mean improves
Frequently asked questions
Is an automated judge required?
No. It can speed up triage if its rubric is validated; ambiguous or costly cases still need business adjudication.
Is a mean improvement enough?
No. Compare relevant cases and segments, then apply the criteria established before testing.
How many cases are needed?
Volume depends on diversity and risk. Report count, provenance and limitations; no small set establishes general safety.
Official references
References consulted on 2 October 2026. The method and worksheet propose checks to adapt to your context; they do not constitute certification.






