The OCR field, narrowed to the models worth testing.

Scan the contenders in seconds. Then choose by the failure your documents cannot afford.

Landscape scanner + Change monitorSee which models deserve a first look.

Which OCR models deserve a first look?

Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.

Do this first

Use the shortlist to choose two contenders, then test the failure your workflow can least afford on your own pages.

Keep this boundary in view

It does not establish one universal winner or accuracy on documents outside the measured slices.

Inspect the evidence and what would change the conclusion
Evidence layer
Reported benchmark survey plus Instavar full-50 workflow run
Observed
March 2026 workflow update
Sample
Four operational workflows on 50 pages, with public benchmark values kept separate

What supports it

The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.

Change our mind if

A different candidate repeatedly avoids the reader's costly failure on a representative local sample.

Claim OCR-SHORTLIST-2026-03 / v2, checked 6 Aug 2026

Scope, evidence, and what this does not prove

What we found

  • The top reported OCR models are now close enough on headline benchmarks that production fit matters more than tiny score gaps.
  • GLM-OCR and PaddleOCR-VL-1.5 still belong in the reported OmniDocBench shortlist.
  • Our hands-on read is more practical: Hunyuan is strongest when coordinates matter, DeepSeek helps when blank-page handling matters, FireRed is the best balanced operational choice, and GLM remains the fastest normal-case workflow.
  • dots.ocr-1.5 belongs in the OCR plus broader visual parsing lane, not as the default scanned-PDF model.
  • Use this page to build the first shortlist, then run a fixed page-type bake-off before rollout.
Update (Mar 2026):
The public shortlist should now be read with a second layer in mind: our newer full-50 workflow benchmark across Hunyuan, DeepSeek, GLM, and FireRed.
That benchmark does not replace the public leaderboard tables below, but it does change the deployment readout: Hunyuan leads on grounded output, DeepSeek is now the second grounded workflow and the strongest blank-page detector, FireRed remains the best balanced workflow, and GLM remains the fastest normal-case path.
For the practical routing answer across those workflows plus dots.ocr-1.5 and PaddleOCR-VL-1.5, see: https://instavar.com/research/ocr/choose-an-ocr-model-by-workflow.

How to use this page

This is a market map and shortlist builder, not the final deployment rule.

The decision path is:

  1. use reported benchmarks to remove weak candidates
  2. use the shortlist table below to pick the first models to test
  3. read the benchmark landscape to understand what each public score does and does not prove
  4. use the workflow-fit guide when the decision depends on page type
  5. use the production checklist before promoting any model

For the scan-heavy benchmarking method behind the practical routing advice, see:

One ranking cannot answer four different risks.

Follow the evidence, then choose the failure your workflow cannot afford.

Measured hereFour operational workflows across one 50-page run, including three blank pages.

Reported elsewhereOmniDocBench scores published by each model author, not rerun on one shared system.

Not establishedAccuracy on your documents, universal model quality, or the quality of every returned visual anchor.

Reported benchmark score

FireRed92.94
GLM94.62
Hunyuan94.10
DeepSeek91.09

Author-reported OmniDocBench results from model papers and repositories.

Reported benchmark score

FireRed92.94
GLM94.62
Hunyuan94.10
DeepSeek91.09

Author-reported OmniDocBench results from model papers and repositories.

Choose a first model to test.

This is a four-model measured slice, not the whole OCR market. The cost of a mistake changes which result deserves attention first.

What would hurt your workflow most?

First model to test for this risk

FireRed

FireRed offers the clearest balance of speed, blank-page handling, and clean structured output.

Watch for
Early release; broad replication is still limited
Change this recommendation if
Another workflow produces fewer consequential errors or needs less cleanup on your representative sample.
Speed and grounded output
Slower processing (seconds per page) →More visual anchors →FireRedGLMHunyuanDeepSeek

Point emphasis follows your selected risk. Positions remain the same measured speed and unverified anchor counts.

  1. 01

    FireRed

    Balanced structured Markdown workflow

    Watch for: Early release; broad replication is still limited

    92.94 reported3.328 sec/page2/3 blank48 anchors
  2. 02

    GLM

    Fastest normal-page workflow

    Watch for: Detected none of the three blank pages

    94.62 reported1.252 sec/page0/3 blank57 anchors
  3. 03

    Hunyuan

    Densest grounded output

    Watch for: Slower and usually needs more normalization

    94.10 reported6.884 sec/page2/3 blank1,517 anchors
  4. 04

    DeepSeek

    Strongest blank-page handling

    Watch for: Slowest workflow in the measured group

    91.09 reported17.591 sec/page3/3 blank926 anchors

Reported scores come from model authors and are not measurements from the same runtime as the hands-on workflow results. Changing the risk changes the editorial order, not the underlying measurements.

Local test / OCR-SHORTLIST-2026-03

Test this claim on your documents.

Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.

Follow OCR research via RSS

Preview the experiment brief
Use Codex to test this bounded OCR claim in my local environment.

Source: https://instavar.com/research/ocr/open-document-ocr-models
Claim ID: OCR-SHORTLIST-2026-03
Claim version: v2, checked 6 Aug 2026
Claim: Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.
Observed evidence: The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.
Evidence layer: Reported benchmark survey plus Instavar full-50 workflow run.
Sample: Four operational workflows on 50 pages, with public benchmark values kept separate.
Boundary: It does not establish one universal winner or accuracy on documents outside the measured slices.
Potential falsifier: A different candidate repeatedly avoids the reader's costly failure on a representative local sample.

Safety and privacy:
- Work read-only by default and keep all documents and outputs local.
- Do not upload documents, reveal credentials, or send document contents to an external service.
- Do not install software, download large models, start paid jobs, or change my environment without asking first.
- If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested.

Test method:
1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration.
2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available.
3. Before running anything, state the primary failure we are testing and the measurement that would expose it.
4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured.
5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page.
6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict.
7. Name the smallest follow-up test that could change the verdict.
What did your local test say?

Save a coarse verdict in this browser. Instavar receives no document contents through this control.