The OCR field, narrowed to the models worth testing.
Scan the contenders in seconds. Then choose by the failure your documents cannot afford.
Which OCR models deserve a first look?
Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.
Do this first
Use the shortlist to choose two contenders, then test the failure your workflow can least afford on your own pages.
Keep this boundary in view
It does not establish one universal winner or accuracy on documents outside the measured slices.
Inspect the evidence and what would change the conclusion
- Evidence layer
- Reported benchmark survey plus Instavar full-50 workflow run
- Observed
- March 2026 workflow update
- Sample
- Four operational workflows on 50 pages, with public benchmark values kept separate
What supports it
The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.
Change our mind if
A different candidate repeatedly avoids the reader's costly failure on a representative local sample.
Claim OCR-SHORTLIST-2026-03 / v2, checked 6 Aug 2026
Scope, evidence, and what this does not prove
What we found
- The top reported OCR models are now close enough on headline benchmarks that production fit matters more than tiny score gaps.
GLM-OCRandPaddleOCR-VL-1.5still belong in the reported OmniDocBench shortlist.- Our hands-on read is more practical:
Hunyuanis strongest when coordinates matter,DeepSeekhelps when blank-page handling matters,FireRedis the best balanced operational choice, andGLMremains the fastest normal-case workflow. dots.ocr-1.5belongs in the OCR plus broader visual parsing lane, not as the default scanned-PDF model.- Use this page to build the first shortlist, then run a fixed page-type bake-off before rollout.
Update (Mar 2026):
The public shortlist should now be read with a second layer in mind: our newer full-50 workflow benchmark acrossHunyuan,DeepSeek,GLM, andFireRed.
That benchmark does not replace the public leaderboard tables below, but it does change the deployment readout:Hunyuanleads on grounded output,DeepSeekis now the second grounded workflow and the strongest blank-page detector,FireRedremains the best balanced workflow, andGLMremains the fastest normal-case path.
For the practical routing answer across those workflows plusdots.ocr-1.5andPaddleOCR-VL-1.5, see: https://instavar.com/research/ocr/choose-an-ocr-model-by-workflow.
How to use this page
This is a market map and shortlist builder, not the final deployment rule.
The decision path is:
- use reported benchmarks to remove weak candidates
- use the shortlist table below to pick the first models to test
- read the benchmark landscape to understand what each public score does and does not prove
- use the workflow-fit guide when the decision depends on page type
- use the production checklist before promoting any model
For the scan-heavy benchmarking method behind the practical routing advice, see:
One ranking cannot answer four different risks.
Follow the evidence, then choose the failure your workflow cannot afford.
Measured hereFour operational workflows across one 50-page run, including three blank pages.
Reported elsewhereOmniDocBench scores published by each model author, not rerun on one shared system.
Not establishedAccuracy on your documents, universal model quality, or the quality of every returned visual anchor.
Reported benchmark score
Author-reported OmniDocBench results from model papers and repositories.
Reported benchmark score
Author-reported OmniDocBench results from model papers and repositories.
Choose a first model to test.
This is a four-model measured slice, not the whole OCR market. The cost of a mistake changes which result deserves attention first.
First model to test for this risk
FireRed
FireRed offers the clearest balance of speed, blank-page handling, and clean structured output.
- Watch for
- Early release; broad replication is still limited
- Change this recommendation if
- Another workflow produces fewer consequential errors or needs less cleanup on your representative sample.
Point emphasis follows your selected risk. Positions remain the same measured speed and unverified anchor counts.
- 01
FireRed
Balanced structured Markdown workflow
Watch for: Early release; broad replication is still limited
92.94 reported3.328 sec/page2/3 blank48 anchors - 02
GLM
Fastest normal-page workflow
Watch for: Detected none of the three blank pages
94.62 reported1.252 sec/page0/3 blank57 anchors - 03
Hunyuan
Densest grounded output
Watch for: Slower and usually needs more normalization
94.10 reported6.884 sec/page2/3 blank1,517 anchors - 04
DeepSeek
Strongest blank-page handling
Watch for: Slowest workflow in the measured group
91.09 reported17.591 sec/page3/3 blank926 anchors
Reported scores come from model authors and are not measurements from the same runtime as the hands-on workflow results. Changing the risk changes the editorial order, not the underlying measurements.
1) Start here: which models belong in your first shortlist?
| Your priority | Recommended first model to test | Why | Second model to test |
| Highest headline benchmark performance | GLM-OCR | Top reported OmniDocBench score among current open releases | PaddleOCR-VL-1.5 |
| Highest grounded workflow output | HunyuanOCR | Strongest grounded workflow in the current full-50 hands-on run | DeepSeek-OCR-2 |
| Best blank-page handling in the current hands-on workflow benchmark | DeepSeek-OCR-2 | Only workflow in the four-model full-50 run to detect all 3/3 blank pages | HunyuanOCR |
| OCR plus broader image parsing (SVG, web, scene text) |
Local test / OCR-SHORTLIST-2026-03
Test this claim on your documents.
Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.
Preview the experiment brief
Use Codex to test this bounded OCR claim in my local environment. Source: https://instavar.com/research/ocr/open-document-ocr-models Claim ID: OCR-SHORTLIST-2026-03 Claim version: v2, checked 6 Aug 2026 Claim: Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test. Observed evidence: The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run. Evidence layer: Reported benchmark survey plus Instavar full-50 workflow run. Sample: Four operational workflows on 50 pages, with public benchmark values kept separate. Boundary: It does not establish one universal winner or accuracy on documents outside the measured slices. Potential falsifier: A different candidate repeatedly avoids the reader's costly failure on a representative local sample. Safety and privacy: - Work read-only by default and keep all documents and outputs local. - Do not upload documents, reveal credentials, or send document contents to an external service. - Do not install software, download large models, start paid jobs, or change my environment without asking first. - If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested. Test method: 1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration. 2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available. 3. Before running anything, state the primary failure we are testing and the measurement that would expose it. 4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured. 5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page. 6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict. 7. Name the smallest follow-up test that could change the verdict.