The OCR field, narrowed to the models worth testing.

Compare the evidence. Choose two models to test on your own documents.

Which OCR models deserve a first look?

Choose two contenders from the shortlist. Test them on your own documents, focusing on the errors your workflow can least afford.

Benchmark scores do not establish a universal winner or accuracy on your documents.

Evidence and limitations

How to use the shortlist

Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.

Use the shortlist to choose two contenders, then test the failure your workflow can least afford on your own pages.

Limits

It does not establish one universal winner or accuracy on documents outside the measured slices.

Evidence layer
Reported benchmark survey plus Instavar full-50 workflow run
Observed
March 2026 workflow update
Sample
Four operational workflows on 50 pages, with public benchmark values kept separate

What supports it

The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.

Change our mind if

A different candidate repeatedly avoids the reader's costly failure on a representative local sample.

What we found

  • Local OmniDocBench v1.6 update: OvisOCR2 led our pinned 1,651-page evaluation at 96.522, followed by GLM-OCR at 94.214.
  • Jina OCR v1: our local FastMTP run scored 88.391. Its text and reading-order results were stronger than Qwen3.8-27B, while its table score was lower.
  • Jina completed all 1,651 pages with zero failed records, but 87 pages hit the 4,096-token output limit. That unresolved cap effect is why Jina remains a challenger rather than a promoted default route.
  • The strongest benchmark leaders are close enough that production fit still matters, while the wider field remains far enough apart for OmniDocBench to separate different output contracts.
  • Our scan-heavy workflow evidence remains a separate decision layer: Hunyuan is strongest when coordinates matter, DeepSeek helps when blank-page handling matters, FireRed is the best balanced operational choice, and GLM remains the fastest normal-case workflow.
  • Use this page to form the shortlist, then test the failure modes that matter on your own pages.
Update (September 2026):
We now have local 1,651-page OmniDocBench v1.6 scores for Jina OCR v1, OvisOCR2, Qwen3.8, Gemma 4, TurboOCR and the earlier OCR workflows under one pinned evaluator. The table below is the canonical local comparison.
These are document-parsing quality results, not a throughput leaderboard or proof of production reliability on scan-heavy documents.
Update (March 2026):
The full-50 workflow benchmark remains the practical routing layer across Hunyuan, DeepSeek, GLM, FireRed, and Qianfan.
For the page-type deployment answer, see: https://instavar.com/research/ocr/choose-an-ocr-model-by-workflow.

Our local 1,651-page OmniDocBench v1.6 results

We evaluated these systems against the same frozen 1,651-page OmniDocBench v1.6 truth and evaluator revision 147cd5ac9472002f5751221d390bf00abdbc0d2f. All failed records stay in the denominator as empty predictions. Outputs that stop at a generation limit remain scored and are counted separately where the run recorded them.

Higher Overall, formula CDM and table TEDS are better. Lower text edit and reading-order edit are better. Overall is the mean of text accuracy (1 - text edit), formula CDM and table TEDS, with each component scaled to 100.

ModelOverallText editFormula CDMTable TEDSReading-order edit
OvisOCR296.5220.02670.97420.94820.1104
GLM-OCR94.2140.03900.96090.90450.1411

Claim OCR-SHORTLIST-2026-03 / v2, checked 6 Aug 2026

One ranking cannot answer four different risks.

Follow the evidence, then choose the failure your workflow cannot afford.

Measured hereFour operational workflows across one 50-page run, including three blank pages.

Reported elsewhereOmniDocBench scores published by each model author, not rerun on one shared system.

Not establishedAccuracy on your documents, universal model quality, or the quality of every returned visual anchor.

Reported benchmark score

FireRed92.94
GLM94.62
Hunyuan94.10
DeepSeek91.09

Author-reported OmniDocBench results from model papers and repositories.

Reported benchmark score

FireRed92.94
GLM94.62
Hunyuan94.10
DeepSeek91.09

Author-reported OmniDocBench results from model papers and repositories.

Choose a first model to test.

This is a four-model measured slice, not the whole OCR market. The cost of a mistake changes which result deserves attention first.

What would hurt your workflow most?

First model to test for this risk

FireRed

FireRed offers the clearest balance of speed, blank-page handling, and clean structured output.

Watch for
Early release; broad replication is still limited
Change this recommendation if
Another workflow produces fewer consequential errors or needs less cleanup on your representative sample.
Speed and grounded output
Slower processing (seconds per page) →More visual anchors →FireRedGLMHunyuanDeepSeek

Point emphasis follows your selected risk. Positions remain the same measured speed and unverified anchor counts.

  1. 01

    FireRed

    Balanced structured Markdown workflow

    Watch for: Early release; broad replication is still limited

    92.94 reported3.328 sec/page2/3 blank48 anchors
  2. 02

    GLM

    Fastest normal-page workflow

    Watch for: Detected none of the three blank pages

    94.62 reported1.252 sec/page0/3 blank57 anchors
  3. 03

    Hunyuan

    Densest grounded output

    Watch for: Slower and usually needs more normalization

    94.10 reported6.884 sec/page2/3 blank1,517 anchors
  4. 04

    DeepSeek

    Strongest blank-page handling

    Watch for: Slowest workflow in the measured group

    91.09 reported17.591 sec/page3/3 blank926 anchors

Reported scores come from model authors and are not measurements from the same runtime as the hands-on workflow results. Changing the risk changes the editorial order, not the underlying measurements.

Local test / OCR-SHORTLIST-2026-03

Test this claim on your documents.

Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.

Follow OCR research via RSS

Preview the experiment brief
Use Codex to test this bounded OCR claim in my local environment.

Source: https://instavar.com/research/ocr/open-document-ocr-models
Claim ID: OCR-SHORTLIST-2026-03
Claim version: v2, checked 6 Aug 2026
Claim: Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.
Observed evidence: The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.
Evidence layer: Reported benchmark survey plus Instavar full-50 workflow run.
Sample: Four operational workflows on 50 pages, with public benchmark values kept separate.
Boundary: It does not establish one universal winner or accuracy on documents outside the measured slices.
Potential falsifier: A different candidate repeatedly avoids the reader's costly failure on a representative local sample.

Safety and privacy:
- Work read-only by default and keep all documents and outputs local.
- Do not upload documents, reveal credentials, or send document contents to an external service.
- Do not install software, download large models, start paid jobs, or change my environment without asking first.
- If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested.

Test method:
1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration.
2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available.
3. Before running anything, state the primary failure we are testing and the measurement that would expose it.
4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured.
5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page.
6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict.
7. Name the smallest follow-up test that could change the verdict.
What did your local test say?

Save a coarse verdict in this browser. Instavar receives no document contents through this control.