Resources · 49

Evaluate AI regressions before changing versions

Compare versions on the same tasks, find local losses and make a decision supported by cases rather than an average.

· 3 min

Method diagram: Task → Independent set → Two versions → Segment losses → Decision Method diagram · steps explained in the text

What this guide helps achieve

  • Isolate the tested change
  • Preserve an independent reference set
  • Inspect failures by segment
  • Prepare shutdown or rollback

Quick check

  • Which component changed?
  • Was the test set used to optimise the candidate?
  • Who resolves annotation disagreements?
  • Does an average hide a sensitive error?
  • Can the previous configuration be restored?

Step-by-step method

  1. 01

    Define task and refusals

    Describe expected output, permitted data and costly errors. Separate accuracy, usefulness, appropriate refusal and action effects. NIST’s framework calls for context-appropriate measurement; it does not provide a universal reliability score.

    Deliverable: criteria and decision owner.

  2. 02

    Separate reference and discovery

    Keep a stable set out of optimisation. Add new anonymised feedback cases separately: they reveal blind spots but should not silently change the comparison denominator.

    Deliverable: versioned sets with provenance.

  3. 03

    Document annotation

    Explain what makes an answer acceptable, partial or incorrect. Have a competent reviewer resolve ambiguity and retain reasons for disagreements. A single reference answer may be too restrictive for an open-ended task.

    Deliverable: rubric and adjudicated cases.

  4. 04

    Compare under equal conditions

    Replay identical inputs with recorded versions, parameters, tools and sources. Repeat trials when outputs vary and report variation. Separate model failure, absent data, unavailable tools and a defective test.

    Deliverable: paired outcomes and minimised traces.

  5. 05

    Decide by risk family

    Inspect losses by language, document type and sensitive situation even when the mean improves. Define stop, correction and limited-release criteria before reading results. One severe case can justify withholding a release.

    Deliverable: decision, reservations and exceptions.

  6. 06

    Replay after release

    Prepare a restorable configuration and reachable owner. Compare early feedback with the test set without treating local success as evidence for all uses. Reassess when data or dependencies change.

    Deliverable: monitoring and rollback procedure.

Reusable worksheet

Complete with your authorised observations. These fields are a working template, not observed results.

FieldInformation to record
Case and segmentStable identifier, language, task type
ExpectedAcceptance criterion and reference evidence
Versions A / BOutcome, repetition, failure cause
DecisionAccept, correct or stop; owner and evidence

Worked example

Illustrative situation

Fictional example: a release improves routine answers but drops uncertainty statements when documents conflict.

Decision and expected evidence

The report isolates that subset, retains source evidence and withholds release for that use until a correction is retested.

Distinguish the mechanisms

MechanismPurposeCheck or limitation
Stable setCompare versionsKeep out of optimisation
New casesFind blind spotsReport separately from historical results
In-service observationUnderstand actual useRespect permissions, minimisation and context

Management indicators

IndicatorWhat it measuresFirst action
Paired lossesPreviously accepted cases now failingInspect each sensitive loss
Annotation disagreementCases with differing judgementsClarify the rubric before comparison
Segment coverageSituations actually representedName untested uses

Common pitfalls

  • Optimising on the final test
  • Changing set and version together
  • Treating an automated judge as truth
  • Accepting critical losses because the mean improves

Frequently asked questions

Is an automated judge required?

No. It can speed up triage if its rubric is validated; ambiguous or costly cases still need business adjudication.

Is a mean improvement enough?

No. Compare relevant cases and segments, then apply the criteria established before testing.

How many cases are needed?

Volume depends on diversity and risk. Report count, provenance and limitations; no small set establishes general safety.

Official references

References consulted on 2 October 2026. The method and worksheet propose checks to adapt to your context; they do not constitute certification.