Resources · 102

OCR: validate document reading before automation

Check omissions, numbers and reading order before transcription feeds search or decisions.

· 2 min

Documents and note taking during a discussion Illustration · fictional scene

What this guide helps achieve

  • Separate source types
  • Build a readable reference
  • Test consequential errors
  • Check structure
  • Define human review

Quick check

  • Is a high score enough?
  • Should all numbers be corrected automatically?

Step-by-step method

  1. 01

    Separate source types

    Identify native text, scanned images and mixed documents. Retain page, region and document version; an existing text layer can already be incorrect.

    Deliverable: source and coordinate inventory.

  2. 02

    Build a readable reference

    Include rotated pages, tables, accents and small characters. Have decisive fields transcribed independently of the engine; do not use its output as ground truth.

    Deliverable: checked reference transcription.

  3. 03

    Test consequential errors

    Check amounts, dates, identifiers, signs and separators. Confusing zero with letter O can change reconciliation even when most words are correct.

    Deliverable: field errors and consequences.

  4. 04

    Check structure

    Inspect columns, cells, headings and missing pages. Compare reading after image preparation while retaining the original source and the settings used.

    Deliverable: order and segmentation test.

  5. 05

    Define human review

    Route ambiguous or consequential fields to review. Retain raw value, correction and owner; replay identical examples after engine or processing changes.

    Deliverable: review and regression rules.

What should happen to the reading?

Check the observed state before choosing next steps.

  1. Confirmed reading

    Accept the field and retain its source region.

  2. Ambiguous reading

    Request review before using the field.

  3. Missing page or region

    Recover the source and check coverage.

Reusable worksheet

Complete with your authorised observations. These fields are a working template, not observed results.

FieldInformation to record
FieldRead value, page and region
ReferenceExpected value and reviewer
CorrectionDecision, author and version

Worked example

Illustrative situation

Illustrative case: an identifier in a scan includes O and 0. The transcription reads fluently but produces a different reference.

Decision and expected evidence

Compare the original region, preserve the ambiguity and prevent automated use until confirmed.

Distinguish the mechanisms

MechanismPurposeCheck or limitation
Native textExtraction without recognitionCheck the existing text layer
OCRReading pixelsPreserve links to source images
ReviewCorrecting a fieldTrace source and change

Management indicators

IndicatorWhat it measuresFirst action
Consequential fieldsFields checked against a referenceReview failures
Page coverageExpected and actually read pagesAddress omissions

Common pitfalls

  • Use OCR output as its own ground truth
  • Approve an amount using only overall text quality
  • Lose the link to source page and region

Frequently asked questions

Is a high score enough?

No. Check fields whose errors change the outcome; an overall score can conceal an amount error.

Should all numbers be corrected automatically?

Avoid assumed corrections. Compare the source region and request review when several readings are plausible.

Official references

References consulted: . The method and worksheet propose checks to adapt to your context; they do not constitute certification.