The OCR field, narrowed to the models worth testing.
Compare the evidence. Choose two models to test on your own documents.
Which OCR models deserve a first look?
Choose two contenders from the shortlist. Test them on your own documents, focusing on the errors your workflow can least afford.
Benchmark scores do not establish a universal winner or accuracy on your documents.
Evidence and limitations
How to use the shortlist
Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.
Use the shortlist to choose two contenders, then test the failure your workflow can least afford on your own pages.
Limits
It does not establish one universal winner or accuracy on documents outside the measured slices.
- Evidence layer
- Reported benchmark survey plus Instavar full-50 workflow run
- Observed
- March 2026 workflow update
- Sample
- Four operational workflows on 50 pages, with public benchmark values kept separate
What supports it
The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.
Change our mind if
A different candidate repeatedly avoids the reader's costly failure on a representative local sample.
What we found
- OvisOCR2 update: it led our frozen 118-page transcription audit across ten tested systems. This is a local text-error result, not a reproduction of its authors'
96.58OmniDocBench v1.6 score. See the OvisOCR2 evidence breakdown. - The top reported OCR models are now close enough on headline benchmarks that production fit matters more than tiny score gaps.
GLM-OCRandPaddleOCR-VL-1.5still belong in the reported OmniDocBench shortlist.- Our hands-on read is more practical:
Hunyuanis strongest when coordinates matter,DeepSeekhelps when blank-page handling matters,FireRedis the best balanced operational choice, andGLMremains the fastest normal-case workflow. dots.ocr-1.5belongs in the OCR plus broader visual parsing lane, not as the default scanned-PDF model.- Use this page to build the first shortlist, then run a fixed page-type bake-off before rollout.
Update (Mar 2026):
The public shortlist should now be read with a second layer in mind: our newer full-50 workflow benchmark acrossHunyuan,DeepSeek,GLM, andFireRed.
That benchmark does not replace the public leaderboard tables below, but it does change the deployment readout:Hunyuanleads on grounded output,DeepSeekis now the second grounded workflow and the strongest blank-page detector,FireRedremains the best balanced workflow, andGLMremains the fastest normal-case path.
For the practical routing answer across those workflows plusdots.ocr-1.5andPaddleOCR-VL-1.5, see: https://instavar.com/research/ocr/choose-an-ocr-model-by-workflow.
OvisOCR2: what we measured and what the authors report
OvisOCR2 belongs on a transcription shortlist because of measured local results, not simply because it is a newer release. Keep three different pieces of evidence separate:
| Evidence | Result | What it establishes |
| Authors' model card | 96.58 overall on OmniDocBench v1.6 | A publisher-reported document-parsing score, not our measurement |
| Our frozen 118-page transcription audit | Macro character error rate 0.096742; macro word error rate 0.189194 | Lowest aggregate transcription error among our ten tested systems under this scoring contract |
| Our 1,651-page OmniDocBench v1.6 inference run | 1,651 successful records, zero failed records, 10 generation-limit stops | Full inference coverage; the retained evaluator summary reports an error and no overall or component scores |
The authors describe a 0.8B page parser and report the 96.58 score in the
OvisOCR2 model card. We have not
verified a local score that reproduces that claim. Finishing inference is not the
same as completing evaluation, and a successful record can still contain a
truncated answer.
Our matched transcription comparison
We rescored retained outputs from a 143-page run using one frozen 118-page reference set. We excluded 25 pages because their reference transcriptions were partial, not because a particular model performed badly. Every system received the same remaining pages, normalization, and failure rule.
These are the three lowest macro character-error results in that ten-system comparison. Lower is better; macro means each page has equal weight.
Claim OCR-SHORTLIST-2026-03 / v2, checked 6 Aug 2026
One ranking cannot answer four different risks.
Follow the evidence, then choose the failure your workflow cannot afford.
Measured hereFour operational workflows across one 50-page run, including three blank pages.
Reported elsewhereOmniDocBench scores published by each model author, not rerun on one shared system.
Not establishedAccuracy on your documents, universal model quality, or the quality of every returned visual anchor.
Reported benchmark score
Author-reported OmniDocBench results from model papers and repositories.
Reported benchmark score
Author-reported OmniDocBench results from model papers and repositories.
Choose a first model to test.
This is a four-model measured slice, not the whole OCR market. The cost of a mistake changes which result deserves attention first.
First model to test for this risk
FireRed
FireRed offers the clearest balance of speed, blank-page handling, and clean structured output.
- Watch for
- Early release; broad replication is still limited
- Change this recommendation if
- Another workflow produces fewer consequential errors or needs less cleanup on your representative sample.
Point emphasis follows your selected risk. Positions remain the same measured speed and unverified anchor counts.
- 01
FireRed
Balanced structured Markdown workflow
Watch for: Early release; broad replication is still limited
92.94 reported3.328 sec/page2/3 blank48 anchors - 02
GLM
Fastest normal-page workflow
Watch for: Detected none of the three blank pages
94.62 reported1.252 sec/page0/3 blank57 anchors - 03
Hunyuan
Densest grounded output
Watch for: Slower and usually needs more normalization
94.10 reported6.884 sec/page2/3 blank1,517 anchors - 04
DeepSeek
Strongest blank-page handling
Watch for: Slowest workflow in the measured group
91.09 reported17.591 sec/page3/3 blank926 anchors
Reported scores come from model authors and are not measurements from the same runtime as the hands-on workflow results. Changing the risk changes the editorial order, not the underlying measurements.
How to use this page
This is a market map and shortlist builder, not the final deployment rule.
The decision path is:
- use reported benchmarks to remove weak candidates
- use the shortlist table below to pick the first models to test
- read the benchmark landscape to understand what each public score does and does not prove
- use the workflow-fit guide when the decision depends on page type
- use the production checklist before promoting any model
For the scan-heavy benchmarking method behind the practical routing advice, see: https://instavar.com/research/ocr/benchmarking-ocr-on-scan-heavy-pdfs.
Quick terms used below:
grounded outputmeans the model returns text with page coordinates, so a reviewer can trace the text back to a region on the original page.blank-page handlingmeans the model can return nothing when a scanned page is actually empty, instead of inventing text.OmniDocBenchis a public document OCR benchmark. It is useful for shortlisting, but it is not a substitute for testing your own page types.
1) Start here: which models belong in your first shortlist?
| Your priority | Recommended first model to test | Why |
Local test / OCR-SHORTLIST-2026-03
Test this claim on your documents.
Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.
Preview the experiment brief
Use Codex to test this bounded OCR claim in my local environment. Source: https://instavar.com/research/ocr/open-document-ocr-models Claim ID: OCR-SHORTLIST-2026-03 Claim version: v2, checked 6 Aug 2026 Claim: Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test. Observed evidence: The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run. Evidence layer: Reported benchmark survey plus Instavar full-50 workflow run. Sample: Four operational workflows on 50 pages, with public benchmark values kept separate. Boundary: It does not establish one universal winner or accuracy on documents outside the measured slices. Potential falsifier: A different candidate repeatedly avoids the reader's costly failure on a representative local sample. Safety and privacy: - Work read-only by default and keep all documents and outputs local. - Do not upload documents, reveal credentials, or send document contents to an external service. - Do not install software, download large models, start paid jobs, or change my environment without asking first. - If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested. Test method: 1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration. 2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available. 3. Before running anything, state the primary failure we are testing and the measurement that would expose it. 4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured. 5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page. 6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict. 7. Name the smallest follow-up test that could change the verdict.