Which OCR model should I test first for this workflow?

By Wei Jie Chee · 13 Mar 2026, 00:00 Z

Download printable cheat-sheet (CC-BY 4.0)
Choice-makerMatch the first model to the failure you cannot afford.

OCR selection is a routing decision: blank pages, grounding, speed, tables, and diagrams reward different first tests.

Do this first

Name the failure that would do the most damage, then begin with the route that is strongest against that failure.

Keep this boundary in view

The route is a first-test recommendation, not a production guarantee or a fixed winner across corpora.

Inspect the evidence and what would change the conclusion
Evidence layer
Cross-run workflow synthesis
Observed
Updated through 6 Aug 2026
Sample
Original 31-PDF pilot plus later 50-page operational comparisons

What supports it

The routing guidance combines reported model evidence with Instavar's scan-heavy pilot and later full-50 workflow runs.

Change our mind if

One model meets the reader's quality, latency, and failure constraints across a representative sample without routing.

Claim OCR-ROUTING-2026-03 / v2, checked 6 Aug 2026

This guide answers the deployment question behind most OCR benchmarks: which model should handle which page type.

If you searched for GLM OCR vs Mistral OCR, GLM OCR vs PaddleOCR, PaddleOCR vs Mistral OCR, dots OCR 1.5, or best OCR model 2026, this is the practical routing page. It does not claim one universal winner. It maps each workflow bottleneck to the first model worth testing.

The answer is not one universal winner. It is a routing map. A page of notes, a worksheet with diagrams, a blank scan, and a table-heavy page fail in different ways, so the best workflow sends each page type to the model that handles it best.

The short version

  • HunyuanOCR is the strongest first test when grounded output matters, meaning text needs to stay tied to page coordinates.
  • DeepSeek-OCR-2 is useful when blank-page handling and grounded fallback behavior matter.
  • FireRed-OCR remains the best balanced operational path for clean markdown on text-first pages.
  • GLM-OCR remains valuable when a question depends on a small local visual such as a graph, apparatus, particle diagram, or reaction scheme.
  • Qianfan, dots.ocr-1.5, PaddleOCR-VL-1.5, and managed APIs still belong in the map when their specific workflow constraints match.
Update (Mar 2026):
The newer full-50 workflow benchmark widened the practical ranking beyond the original FireRed versus GLM routing story.
Hunyuan is now the strongest grounded workflow, DeepSeek is the second grounded workflow and the only one to detect all 3/3 blank pages in the current full-50 run, FireRed remains the best balanced workflow, and GLM remains the fastest normal-case workflow.
Qianfan is now a promoted workflow and belongs in the routing map as the markdown-oriented fallback lane.
A page-level router across all five promoted workflows (FireRed, GLM, Hunyuan, DeepSeek, Qianfan) is operational and under active iteration, but not yet promoted as a default - see Section 10 for the early benchmark results.
That means the deployment answer is now a five-lane map, not just a single FireRed/GLM split.

The decision path

Most OCR comparisons start with the benchmark table. We start with page failure modes.

The scan-heavy pilot first made GLM-OCR look like the safest default. That changed after the FireRed-OCR wrapper stopped hallucinating on near-blank pages and preserved page images. The later workflow benchmark then added

Local test / OCR-ROUTING-2026-03

Test this claim on your documents.

Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.

Follow OCR research via RSS

Preview the experiment brief
Use Codex to test this bounded OCR claim in my local environment.

Source: https://instavar.com/research/ocr/choose-an-ocr-model-by-workflow
Claim ID: OCR-ROUTING-2026-03
Claim version: v2, checked 6 Aug 2026
Claim: OCR selection is a routing decision: blank pages, grounding, speed, tables, and diagrams reward different first tests.
Observed evidence: The routing guidance combines reported model evidence with Instavar's scan-heavy pilot and later full-50 workflow runs.
Evidence layer: Cross-run workflow synthesis.
Sample: Original 31-PDF pilot plus later 50-page operational comparisons.
Boundary: The route is a first-test recommendation, not a production guarantee or a fixed winner across corpora.
Potential falsifier: One model meets the reader's quality, latency, and failure constraints across a representative sample without routing.

Safety and privacy:
- Work read-only by default and keep all documents and outputs local.
- Do not upload documents, reveal credentials, or send document contents to an external service.
- Do not install software, download large models, start paid jobs, or change my environment without asking first.
- If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested.

Test method:
1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration.
2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available.
3. Before running anything, state the primary failure we are testing and the measurement that would expose it.
4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured.
5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page.
6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict.
7. Name the smallest follow-up test that could change the verdict.
What did your local test say?

Save a coarse verdict in this browser. Instavar receives no document contents through this control.