How should difficult scanned PDF pages be repaired and routed?

By Wei Jie Chee · 28 Mar 2026, 00:00 Z

Download printable cheat-sheet (CC-BY 4.0)
Failure owner + One-off converterPreserve good text layers and route difficult pages.

Preserve trustworthy text layers, classify the failure, and route image-only or damaged pages instead of reprocessing every PDF identically.

Do this first

Inspect the existing text layer first, then route only the pages that are blank, damaged, or structurally unreliable.

Keep this boundary in view

The five-model slice does not cover every scanner, language, archive rule, or document management system.

Inspect the evidence and what would change the conclusion
Evidence layer
Five-model scan-heavy comparison
Observed
March 2026
Sample
50 scanned pages across seven page types

What supports it

The practical model comparisons use a 50-page scan-heavy slice across seven page types, alongside workflow guidance for existing text layers.

Change our mind if

A single fixed OCR path preserves good text layers and meets the reader's accuracy and throughput constraints across the full local mix.

Claim OCR-SCANNED-PDF-50-V1 / v1, checked 6 Aug 2026

This guide answers a practical build question: if your PDFs are scanned images, which OCR model should handle each kind of page.

The short version

  • No single model won all document types in the 50-page benchmark.
  • Qianfan had the lowest aggregate character error rate, or CER, at 12.8%.
  • The more useful result is by page type: GLM dominated diagram pages at 6.1% CER, Hunyuan was strong on low-contrast scans at 6.6% CER, and Qianfan swept text, tables, formulas, and worksheets.
  • The right production answer is a routing rule, not one default model for every scanned PDF.

The one-minute decision path

Scanned PDFs fail differently from born-digital PDFs because the text is inside page images. A clean notes page, a diagram question, a faint scan, and a blank separator page need different safeguards.

Read this page in three passes:

  1. use the quick routing table below for the first implementation choice
  2. check the page-type results before trusting the aggregate score
  3. use the failure notes to decide where to add fallback models or human review
If the scanned page is...Start with...Why
text-first notes, tables, formulas, or worksheetsQianfanlowest measured error across those page types
diagram-heavy or figure-linkedGLMstrongest measured diagram-page result
low-contrast or faintQianfan, with Hunyuan as fallback

Local test / OCR-SCANNED-PDF-50-V1

Test this claim on your documents.

Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.

Follow OCR research via RSS

Preview the experiment brief
Use Codex to test this bounded OCR claim in my local environment.

Source: https://instavar.com/research/ocr/ocr-for-scanned-pdfs
Claim ID: OCR-SCANNED-PDF-50-V1
Claim version: v1, checked 6 Aug 2026
Claim: Preserve trustworthy text layers, classify the failure, and route image-only or damaged pages instead of reprocessing every PDF identically.
Observed evidence: The practical model comparisons use a 50-page scan-heavy slice across seven page types, alongside workflow guidance for existing text layers.
Evidence layer: Five-model scan-heavy comparison.
Sample: 50 scanned pages across seven page types.
Boundary: The five-model slice does not cover every scanner, language, archive rule, or document management system.
Potential falsifier: A single fixed OCR path preserves good text layers and meets the reader's accuracy and throughput constraints across the full local mix.

Safety and privacy:
- Work read-only by default and keep all documents and outputs local.
- Do not upload documents, reveal credentials, or send document contents to an external service.
- Do not install software, download large models, start paid jobs, or change my environment without asking first.
- If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested.

Test method:
1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration.
2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available.
3. Before running anything, state the primary failure we are testing and the measurement that would expose it.
4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured.
5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page.
6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict.
7. Name the smallest follow-up test that could change the verdict.
What did your local test say?

Save a coarse verdict in this browser. Instavar receives no document contents through this control.