The OCR field, narrowed to the models worth testing.

Compare the evidence. Choose two models to test on your own documents.

Which OCR models deserve a first look?

Choose two contenders from the shortlist. Test them on your own documents, focusing on the errors your workflow can least afford.

Benchmark scores do not establish a universal winner or accuracy on your documents.

Evidence and limitations

How to use the shortlist

Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.

Use the shortlist to choose two contenders, then test the failure your workflow can least afford on your own pages.

Limits

It does not establish one universal winner or accuracy on documents outside the measured slices.

Evidence layer
Reported benchmark survey plus Instavar full-50 workflow run
Observed
March 2026 workflow update
Sample
Four operational workflows on 50 pages, with public benchmark values kept separate

What supports it

The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.

Change our mind if

A different candidate repeatedly avoids the reader's costly failure on a representative local sample.

What we found

  • OvisOCR2 update: it led our frozen 118-page transcription audit across ten tested systems. This is a local text-error result, not a reproduction of its authors' 96.58 OmniDocBench v1.6 score. See the OvisOCR2 evidence breakdown.
  • The top reported OCR models are now close enough on headline benchmarks that production fit matters more than tiny score gaps.
  • GLM-OCR and PaddleOCR-VL-1.5 still belong in the reported OmniDocBench shortlist.
  • Our hands-on read is more practical: Hunyuan is strongest when coordinates matter, DeepSeek helps when blank-page handling matters, FireRed is the best balanced operational choice, and GLM remains the fastest normal-case workflow.
  • dots.ocr-1.5 belongs in the OCR plus broader visual parsing lane, not as the default scanned-PDF model.
  • Use this page to build the first shortlist, then run a fixed page-type bake-off before rollout.
Update (Mar 2026):
The public shortlist should now be read with a second layer in mind: our newer full-50 workflow benchmark across Hunyuan, DeepSeek, GLM, and FireRed.
That benchmark does not replace the public leaderboard tables below, but it does change the deployment readout: Hunyuan leads on grounded output, DeepSeek is now the second grounded workflow and the strongest blank-page detector, FireRed remains the best balanced workflow, and GLM remains the fastest normal-case path.
For the practical routing answer across those workflows plus dots.ocr-1.5 and PaddleOCR-VL-1.5, see: https://instavar.com/research/ocr/choose-an-ocr-model-by-workflow.

OvisOCR2: what we measured and what the authors report

OvisOCR2 belongs on a transcription shortlist because of measured local results, not simply because it is a newer release. Keep three different pieces of evidence separate:

EvidenceResultWhat it establishes
Authors' model card96.58 overall on OmniDocBench v1.6A publisher-reported document-parsing score, not our measurement
Our frozen 118-page transcription auditMacro character error rate 0.096742; macro word error rate 0.189194Lowest aggregate transcription error among our ten tested systems under this scoring contract
Our 1,651-page OmniDocBench v1.6 inference run1,651 successful records, zero failed records, 10 generation-limit stopsFull inference coverage; the retained evaluator summary reports an error and no overall or component scores

The authors describe a 0.8B page parser and report the 96.58 score in the OvisOCR2 model card. We have not verified a local score that reproduces that claim. Finishing inference is not the same as completing evaluation, and a successful record can still contain a truncated answer.

Our matched transcription comparison

We rescored retained outputs from a 143-page run using one frozen 118-page reference set. We excluded 25 pages because their reference transcriptions were partial, not because a particular model performed badly. Every system received the same remaining pages, normalization, and failure rule.

These are the three lowest macro character-error results in that ten-system comparison. Lower is better; macro means each page has equal weight.

Claim OCR-SHORTLIST-2026-03 / v2, checked 6 Aug 2026

One ranking cannot answer four different risks.

Follow the evidence, then choose the failure your workflow cannot afford.

Measured hereFour operational workflows across one 50-page run, including three blank pages.

Reported elsewhereOmniDocBench scores published by each model author, not rerun on one shared system.

Not establishedAccuracy on your documents, universal model quality, or the quality of every returned visual anchor.

Reported benchmark score

FireRed92.94
GLM94.62
Hunyuan94.10
DeepSeek91.09

Author-reported OmniDocBench results from model papers and repositories.

Reported benchmark score

FireRed92.94
GLM94.62
Hunyuan94.10
DeepSeek91.09

Author-reported OmniDocBench results from model papers and repositories.

Choose a first model to test.

This is a four-model measured slice, not the whole OCR market. The cost of a mistake changes which result deserves attention first.

What would hurt your workflow most?

First model to test for this risk

FireRed

FireRed offers the clearest balance of speed, blank-page handling, and clean structured output.

Watch for
Early release; broad replication is still limited
Change this recommendation if
Another workflow produces fewer consequential errors or needs less cleanup on your representative sample.
Speed and grounded output
Slower processing (seconds per page) →More visual anchors →FireRedGLMHunyuanDeepSeek

Point emphasis follows your selected risk. Positions remain the same measured speed and unverified anchor counts.

  1. 01

    FireRed

    Balanced structured Markdown workflow

    Watch for: Early release; broad replication is still limited

    92.94 reported3.328 sec/page2/3 blank48 anchors
  2. 02

    GLM

    Fastest normal-page workflow

    Watch for: Detected none of the three blank pages

    94.62 reported1.252 sec/page0/3 blank57 anchors
  3. 03

    Hunyuan

    Densest grounded output

    Watch for: Slower and usually needs more normalization

    94.10 reported6.884 sec/page2/3 blank1,517 anchors
  4. 04

    DeepSeek

    Strongest blank-page handling

    Watch for: Slowest workflow in the measured group

    91.09 reported17.591 sec/page3/3 blank926 anchors

Reported scores come from model authors and are not measurements from the same runtime as the hands-on workflow results. Changing the risk changes the editorial order, not the underlying measurements.

Local test / OCR-SHORTLIST-2026-03

Test this claim on your documents.

Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.

Follow OCR research via RSS

Preview the experiment brief
Use Codex to test this bounded OCR claim in my local environment.

Source: https://instavar.com/research/ocr/open-document-ocr-models
Claim ID: OCR-SHORTLIST-2026-03
Claim version: v2, checked 6 Aug 2026
Claim: Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.
Observed evidence: The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.
Evidence layer: Reported benchmark survey plus Instavar full-50 workflow run.
Sample: Four operational workflows on 50 pages, with public benchmark values kept separate.
Boundary: It does not establish one universal winner or accuracy on documents outside the measured slices.
Potential falsifier: A different candidate repeatedly avoids the reader's costly failure on a representative local sample.

Safety and privacy:
- Work read-only by default and keep all documents and outputs local.
- Do not upload documents, reveal credentials, or send document contents to an external service.
- Do not install software, download large models, start paid jobs, or change my environment without asking first.
- If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested.

Test method:
1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration.
2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available.
3. Before running anything, state the primary failure we are testing and the measurement that would expose it.
4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured.
5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page.
6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict.
7. Name the smallest follow-up test that could change the verdict.
What did your local test say?

Save a coarse verdict in this browser. Instavar receives no document contents through this control.