Resources · 102
OCR: validate document reading before automation
Check omissions, numbers and reading order before transcription feeds search or decisions.
· 2 min
What this guide helps achieve
- Separate source types
- Build a readable reference
- Test consequential errors
- Check structure
- Define human review
Quick check
- Is a high score enough?
- Should all numbers be corrected automatically?
Step-by-step method
- 01
Separate source types
Identify native text, scanned images and mixed documents. Retain page, region and document version; an existing text layer can already be incorrect.
Deliverable: source and coordinate inventory.
- 02
Build a readable reference
Include rotated pages, tables, accents and small characters. Have decisive fields transcribed independently of the engine; do not use its output as ground truth.
Deliverable: checked reference transcription.
- 03
Test consequential errors
Check amounts, dates, identifiers, signs and separators. Confusing zero with letter O can change reconciliation even when most words are correct.
Deliverable: field errors and consequences.
- 04
Check structure
Inspect columns, cells, headings and missing pages. Compare reading after image preparation while retaining the original source and the settings used.
Deliverable: order and segmentation test.
- 05
Define human review
Route ambiguous or consequential fields to review. Retain raw value, correction and owner; replay identical examples after engine or processing changes.
Deliverable: review and regression rules.
What should happen to the reading?
Check the observed state before choosing next steps.
Confirmed reading
Accept the field and retain its source region.
Ambiguous reading
Request review before using the field.
Missing page or region
Recover the source and check coverage.
Reusable worksheet
Complete with your authorised observations. These fields are a working template, not observed results.
| Field | Information to record |
|---|---|
| Field | Read value, page and region |
| Reference | Expected value and reviewer |
| Correction | Decision, author and version |
Worked example
Illustrative situation
Illustrative case: an identifier in a scan includes O and 0. The transcription reads fluently but produces a different reference.
Decision and expected evidence
Compare the original region, preserve the ambiguity and prevent automated use until confirmed.
Distinguish the mechanisms
| Mechanism | Purpose | Check or limitation |
|---|---|---|
| Native text | Extraction without recognition | Check the existing text layer |
| OCR | Reading pixels | Preserve links to source images |
| Review | Correcting a field | Trace source and change |
Management indicators
| Indicator | What it measures | First action |
|---|---|---|
| Consequential fields | Fields checked against a reference | Review failures |
| Page coverage | Expected and actually read pages | Address omissions |
Common pitfalls
- Use OCR output as its own ground truth
- Approve an amount using only overall text quality
- Lose the link to source page and region
Frequently asked questions
Is a high score enough?
No. Check fields whose errors change the outcome; an overall score can conceal an amount error.
Should all numbers be corrected automatically?
Avoid assumed corrections. Compare the source region and request review when several readings are plausible.
Official references
References consulted: . The method and worksheet propose checks to adapt to your context; they do not constitute certification.






