Essential cookies keep Instavar working. Optional analytics help us understand how the site is used and link your first-visit source to your Studio account after sign-in, for up to 180 days. Cookie Policy
Manage Cookie Preferences
Service reliability telemetry, including Sentry error monitoring and Vercel Speed Insights, stays enabled so we can secure the product and diagnose failures.
Choose two contenders from the shortlist. Test them on your own documents, focusing on the errors your workflow can least afford.
Benchmark scores do not establish a universal winner or accuracy on your documents.
Evidence and limitations
How to use the shortlist
Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.
Use the shortlist to choose two contenders, then test the failure your workflow can least afford on your own pages.
Limits
It does not establish one universal winner or accuracy on documents outside the measured slices.
Evidence layer
Reported benchmark survey plus Instavar full-50 workflow run
Observed
March 2026 workflow update
Sample
Four operational workflows on 50 pages, with public benchmark values kept separate
What supports it
The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.
Change our mind if
A different candidate repeatedly avoids the reader's costly failure on a representative local sample.
What we found
Local OmniDocBench v1.6 update:OvisOCR2 led our pinned 1,651-page evaluation at 96.522, followed by GLM-OCR at 94.214.
Jina OCR v1: our local FastMTP run scored 88.391. Its text and reading-order results were stronger than Qwen3.8-27B, while its table score was lower.
Jina completed all 1,651 pages with zero failed records, but 87 pages hit the 4,096-token output limit. That unresolved cap effect is why Jina remains a challenger rather than a promoted default route.
The strongest benchmark leaders are close enough that production fit still matters, while the wider field remains far enough apart for OmniDocBench to separate different output contracts.
Our scan-heavy workflow evidence remains a separate decision layer: Hunyuan is strongest when coordinates matter, DeepSeek helps when blank-page handling matters, FireRed is the best balanced operational choice, and GLM remains the fastest normal-case workflow.
Use this page to form the shortlist, then test the failure modes that matter on your own pages.
Update (September 2026): We now have local 1,651-page OmniDocBench v1.6 scores for Jina OCR v1, OvisOCR2, Qwen3.8, Gemma 4, TurboOCR and the earlier OCR workflows under one pinned evaluator. The table below is the canonical local comparison. These are document-parsing quality results, not a throughput leaderboard or proof of production reliability on scan-heavy documents.
We evaluated these systems against the same frozen 1,651-page OmniDocBench v1.6 truth and evaluator revision 147cd5ac9472002f5751221d390bf00abdbc0d2f. All failed records stay in the denominator as empty predictions. Outputs that stop at a generation limit remain scored and are counted separately where the run recorded them.
Higher Overall, formula CDM and table TEDS are better. Lower text edit and reading-order edit are better. Overall is the mean of text accuracy (1 - text edit), formula CDM and table TEDS, with each component scaled to 100.
Model
Overall
Text edit
Formula CDM
Table TEDS
Reading-order edit
OvisOCR2
96.522
0.0267
0.9742
0.9482
0.1104
GLM-OCR
94.214
0.0390
0.9609
0.9045
0.1411
Claim OCR-SHORTLIST-2026-03 / v2, checked 6 Aug 2026
One ranking cannot answer four different risks.
Follow the evidence, then choose the failure your workflow cannot afford.
Measured hereFour operational workflows across one 50-page run, including three blank pages.
Reported elsewhereOmniDocBench scores published by each model author, not rerun on one shared system.
Not establishedAccuracy on your documents, universal model quality, or the quality of every returned visual anchor.
Reported benchmark score
FireRed92.94
GLM94.62
Hunyuan94.10
DeepSeek91.09
Author-reported OmniDocBench results from model papers and repositories.
Reported benchmark score
FireRed92.94
GLM94.62
Hunyuan94.10
DeepSeek91.09
Author-reported OmniDocBench results from model papers and repositories.
Choose a first model to test.
This is a four-model measured slice, not the whole OCR market. The cost of a mistake changes which result deserves attention first.
First model to test for this risk
FireRed
FireRed offers the clearest balance of speed, blank-page handling, and clean structured output.
Watch for
Early release; broad replication is still limited
Change this recommendation if
Another workflow produces fewer consequential errors or needs less cleanup on your representative sample.
Speed and grounded output
Point emphasis follows your selected risk. Positions remain the same measured speed and unverified anchor counts.
01
FireRed
Balanced structured Markdown workflow
Watch for: Early release; broad replication is still limited
92.94 reported3.328 sec/page2/3 blank48 anchors
02
GLM
Fastest normal-page workflow
Watch for: Detected none of the three blank pages
94.62 reported1.252 sec/page0/3 blank57 anchors
03
Hunyuan
Densest grounded output
Watch for: Slower and usually needs more normalization
Reported scores come from model authors and are not measurements from the same runtime as the hands-on workflow results. Changing the risk changes the editorial order, not the underlying measurements.
OvisOCR2: what we measured and what the authors report
OvisOCR2 belongs on a transcription shortlist because of measured local results,
not simply because it is a newer release. Keep three different pieces of evidence
separate:
A near reproduction under our local model, prompt, runtime and evaluator contract
Our frozen 118-page transcription audit
Macro character error rate 0.096742; macro word error rate 0.189194
Lowest aggregate transcription error among our ten tested systems under this separate scoring contract
The authors describe a 0.8B page parser and report 96.58 in the
OvisOCR2 model card. Our local
96.522 result is 0.058 points lower. That is strong evidence that the pinned
local route performs similarly on this benchmark, but it does not establish the
same result on other hardware, runtimes, prompts or document collections.
Local test / OCR-SHORTLIST-2026-03
Test this claim on your documents.
Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.
Use Codex to test this bounded OCR claim in my local environment.
Source: https://instavar.com/research/ocr/open-document-ocr-models
Claim ID: OCR-SHORTLIST-2026-03
Claim version: v2, checked 6 Aug 2026
Claim: Public scores can identify contenders, but the failure a workflow cannot afford should determine the first local test.
Observed evidence: The page separates author-reported public scores from one Instavar four-workflow, 50-page operational run.
Evidence layer: Reported benchmark survey plus Instavar full-50 workflow run.
Sample: Four operational workflows on 50 pages, with public benchmark values kept separate.
Boundary: It does not establish one universal winner or accuracy on documents outside the measured slices.
Potential falsifier: A different candidate repeatedly avoids the reader's costly failure on a representative local sample.
Safety and privacy:
- Work read-only by default and keep all documents and outputs local.
- Do not upload documents, reveal credentials, or send document contents to an external service.
- Do not install software, download large models, start paid jobs, or change my environment without asking first.
- If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested.
Test method:
1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration.
2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available.
3. Before running anything, state the primary failure we are testing and the measurement that would expose it.
4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured.
5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page.
6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict.
7. Name the smallest follow-up test that could change the verdict.
FireRed-OCR
90.508
0.0605
0.9470
0.8287
0.1544
Qianfan-OCR
89.653
0.0983
0.9425
0.8454
0.1727
HunyuanOCR document parsing
89.014
0.0759
0.8730
0.8733
0.1706
Qwen3.8-27B
88.516
0.1164
0.9179
0.8540
0.1854
Jina-OCR-v1
88.391
0.0784
0.9176
0.8126
0.1709
Gemma 4
66.663
0.3637
0.7558
0.6078
0.3516
DeepSeek-OCR
66.291
0.3314
0.7391
0.5810
0.3524
Chandra-OCR
66.014
0.0616
0.9029
0.1392
0.1650
HunyuanOCR spotting
51.589
0.2760
0.8150
0.0087
0.3379
TurboOCR Tiny
34.970
0.3096
0.3587
0.0000
0.3810
The comparison is like-for-like at the truth and evaluator boundary, but the inference provenance is not identical. OvisOCR2, Jina, Qwen, Gemma and TurboOCR are fresh full-suite runs. The Hunyuan document-parsing row is also a full-suite run, using its explicitly named adaptive repetition fallback. GLM, FireRed, Qianfan, DeepSeek, Chandra and Hunyuan spotting reused 1,355 v1.5 predictions only where decoded page pixels matched exactly, then freshly inferred the 296 changed or new v1.6 pages before all 1,651 predictions were rescored.
The output contracts also differ. TurboOCR is text-only, while Hunyuan spotting is a coordinate-oriented OCR route rather than a document parser. Their low table scores should not be read as evidence that their narrower text or grounding tasks failed. The table is therefore a document-parsing comparison, not one universal ranking of every OCR capability.
What Jina OCR v1 adds
Jina finished only 0.125 Overall points behind Qwen3.8-27B, but the two systems reached that result differently. Jina had better text edit (0.0784 versus 0.1164) and reading-order edit (0.1709 versus 0.1854). Qwen had better table TEDS (0.8540 versus 0.8126). Formula CDM was effectively tied.
The local Jina run used vLLM 0.22.1, BF16, FastMTP with three speculative tokens, batch size 8, and a 4,096-token output ceiling on an RTX 3090 Ti. The model load occupied about 6.31 GiB. It completed every page, but 87/1,651 pages, about 5.3%, reached the output limit.
Jina's model card reports 91.14 on OmniDocBench v1.6. Our local 88.391 is 2.749 points lower. The output ceiling is a plausible contributor, but this run does not establish causality. The smallest useful follow-up is to freeze those 87 pages, rerun only them with an 8,192-token ceiling, and rescore the baseline and repair separately.
Jina-OCR-v1 is released under CC-BY-NC-4.0. Treat the current local route as a research and evaluation lane unless commercial licensing is resolved.
Our matched transcription comparison
We rescored retained outputs from a 143-page run using one frozen 118-page
reference set. We excluded 25 pages because their reference transcriptions were
partial, not because a particular model performed badly. Every system received
the same remaining pages, normalization, and failure rule.
These are the three lowest macro character-error results in that ten-system
comparison. Lower is better; macro means each page has equal weight.
System
Macro character error rate
Macro word error rate
Unique page wins by character error
OvisOCR2
0.096742
0.189194
24
FireRed-OCR
0.127700
0.234079
14
Chandra-OCR-2
0.129275
0.214995
34
OvisOCR2 also led the median and character-weighted or word-weighted aggregate
error measures. Chandra nevertheless won more individual pages outright. The
practical conclusion is to test OvisOCR2 for transcription, not to assume it wins
on every document type. Character error rate counts substitutions, deletions,
and insertions against the reference; it is not a percentage of pages that are
correct.
The scorer case-folds text, normalizes Unicode and whitespace, and removes HTML,
image links, and lightweight markup. It caps each page at 12,000 characters and
3,000 words; eight OvisOCR2 pages were scorer-capped. Missing or unusable outputs
remain in the denominator as empty predictions. OvisOCR2 had no empty outputs
in this audit. These scoring caps are separate from the generation-limit stops
in the 1,651-page run.
This audit measures normalized transcription. It does not validate merged
table cells, formula structure, layout coordinates, or whether generated image
references resolve to usable crops. Our retained OvisOCR2 workflow emitted image
tags without the corresponding retained crops, so usable visual-region export
needs a separate check.
What to test before adopting it
The tested checkpoint was ATH-MaaS/OvisOCR2 revision
65c619d374b55d4152e85150fc1b003700bc1f0c. The earlier local comparison used
vLLM 0.22.1, BF16, the publisher's parsing prompt, and a 16,384-output-token
ceiling. Its observed peak GPU use was about 20.4 GiB on our 24 GB RTX 3090 Ti.
That is an observation for this configuration, not a minimum-memory guarantee.
Before deployment, inspect numeric tables, dense fine print, formulas, and
image-region exports on your own pages. The local 96.522 score is valid for the
pinned benchmark contract, but our retained Ovis output still includes image tags
without the corresponding exported crops. The older February and March
comparisons below remain separate dated snapshots; their benchmark versions and
workflow metrics must not be blended into the full-suite score.
How to use this page
This is a market map and shortlist builder, not the final deployment rule.
The decision path is:
use reported benchmarks to remove weak candidates
use the shortlist table below to pick the first models to test
read the benchmark landscape to understand what each public score does and does not prove
use the workflow-fit guide when the decision depends on page type
use the production checklist before promoting any model
OCR has converged with compact VLM design, and in some workflows these models reduce or replace parts of multi-stage OCR pipelines.
Benchmarks increasingly reward document-level understanding, not just line-level text extraction.
Open releases now include practical deployment paths (vLLM, SGLang, Hugging Face, and in some cases Ollama), reducing integration friction.
3) What the benchmark evidence says (reported)
Before comparing model cards, keep three filters in mind:
Compare only like-for-like benchmarks.
Treat low-sample live leaderboard results as directional, not final.
Validate on your own corpus before production promotion.
3.0 Why one OCR leaderboard score is not enough
A single aggregate OCR score usually hides the failure that will cost you time in production.
That is the pattern builders keep running into when they compare OCR models outside a neat leaderboard. In one Reddit benchmark discussion, users pushed for cost-per-success, latency, and open-model comparisons rather than only flagship model accuracy. In a PaddleOCR-VL-1.5 discussion, users reported strong benchmark scores but still called out table failures, repetition, CPU slowness, and local hardware questions. In a scanned PDF extraction thread, the pain was not that OCR returned no text. The pain was that tables broke, columns shifted, numbers were misread, and the output still needed manual checking.
Read leaderboard scores as a shortlist signal, then test the workflow dimensions that actually fail:
OCR text accuracy: can it read the characters?
Table extraction: can it preserve rows, columns, merged cells, and numeric alignment?
Key information extraction: can it pull fields without moving values into the wrong label?
Visual QA: can it answer questions about diagrams, figures, stamps, signatures, and local images?
Long-document handling: does quality hold after many pages, not just one demo page?
Latency and cost: does the model still make sense at 1,000 pages?
Field-level reliability: can a reviewer trace uncertain values back to the page before they enter a downstream system?
That is why this page is a leaderboard and not a final answer. Use it to build the shortlist, then use the workflow guide and scanned-PDF guide to decide which model should handle each page type.
3.1 OmniDocBench snapshot
The table below consolidates reported OmniDocBench scores from model papers/cards, using v1.5 where explicitly stated.
Model
Params
OmniDocBench (reported)
Notes
Source
GLM-OCR
0.9B
94.62
Strong all-round reported score; very recent release
The top reported scores are now close enough that cost, failure mode, and licensing often matter more than a small benchmark gap.
3.5 FireRed-OCR early evidence snapshot
The FireRed-OCR launch matters because it includes both a technical paper and a benchmark framing centered on structural integrity rather than only text recognition.
FireRed-OCR is not the overall reported OmniDocBench leader, but it is now one of the clearest structure-first challengers in the open OCR field.
If your bottleneck is malformed Markdown or broken document syntax rather than pure text recognition, it belongs in the evaluation lane immediately.
3.6 What hands-on evaluation changed
Public benchmark tables are useful, but real scanned documents can still reorder the shortlist once page type and wrapper quality enter the picture. In our scan-heavy pilot, that is exactly what happened. This page should stay the market map and shortlist, not the final routing answer. For the routing rule and the underlying evidence, use:
3.7 What the newer four-model workflow benchmark changed
The newer full-50 workflow benchmark adds a second layer on top of the public leaderboard story because it compares real operational entrypoints rather than just reported paper/model-card numbers.
Workflow
Mean sec/page
Blank pages detected
Total visual anchors
Practical readout
FireRed
3.328
2/3
48
Best balanced workflow
GLM
1.252
0/3
57
Fastest normal-case workflow
Hunyuan
6.884
2/3
1517
Strongest grounded workflow
DeepSeek
17.591
3/3
Interpretation:
Hunyuan now has the strongest practical case when grounded structure matters more than speed.
DeepSeek is no longer just a markdown-oriented curiosity. It is now the second grounded workflow in the measured stack, although it is also the slowest.
FireRed remains the best balanced operational choice when you want a cleaner markdown-oriented workflow.
GLM remains the fastest typical path, but it is still weak on blank-page handling.
3.2 OlmOCR-Bench snapshot
Reported from LightOnOCR-2 benchmarking (headers/footers excluded setting):
The Elo setup is judged by Gemini 3 Flash in the authors' pipeline and is not a drop-in replacement for independent leaderboard results.
4) Model fit by use case
4.1 Use-case fit matrix
Model
Choose first when
Why it wins there
Watch-outs
HunyuanOCR
You need dense grounded output for extraction or audit-heavy workflows
Strongest grounded workflow in the current full-50 hands-on benchmark
Slower than GLM or FireRed; raw output usually needs more normalization
DeepSeek-OCR-2
You need stronger grounding than GLM or FireRed plus strict blank-page handling
Second grounded workflow in the current hands-on benchmark and the only one to detect 3/3 blank pages
Slowest current workflow; current helper job adds startup overhead
GLM-OCR
You need a strong default baseline across mixed documents
Top-tier reported OmniDocBench result in compact size; multiple serving paths
Very new release; long-tail behavior still needs broad replication
dots.ocr-1.5
You need one model for OCR plus web/screen/scene/SVG parsing
Broad task coverage in a single 3B model family and strong reported release benchmarks
Many benchmark claims are currently model-card/repo reported for this version
FireRed-OCR
You need stricter structural Markdown behavior with formulas and tables
Public training story explicitly targets structural hallucination and syntactic validity
Early-cycle release; benchmark evidence is still author-reported and needs broad replication
DeepSeek-OCR-2
You need markdown-oriented output and mode switching
Reading-order-focused design and dual extraction modes (Free OCR and structured conversion)
Validate complex tables and multilingual edge cases on your own corpus
LightOnOCR-2-1B
You process high page volume and care about cost per page
Strong reported OlmOCR-Bench + throughput profile at 1B scale
Check performance on your language/script distribution
GutenOCR
You need reliable text-to-location grounding for downstream extraction
Grounded OCR is core design objective and first-class output
Weight license is CC-BY-NC; commercial use may be constrained
HunyuanOCR
You want one compact model for broad document tasks
Strong reported compact-model results across parsing-oriented tasks
Custom community license requires legal/compliance review
PaddleOCR-VL-1.5
Your inputs are messy scans/photos and you already run Paddle tooling
Near-frontier reported OmniDocBench score with robustness framing
Confirm accuracy on your distortion mix and template families
4.2 Adoption and maturity signals (Feb 13, 2026 snapshot)
These are not quality scores. They are practical signals for implementation confidence and ecosystem support.
Model
Maturity signal
What it means for rollout
GLM-OCR
Rapid early GitHub/HF uptake after launch
Fast-moving ecosystem, but still early for stability assumptions
dots.ocr-1.5
Fresh Feb 16, 2026 release with expanded task scope
High upside for multi-task use cases, but treat current results as early-cycle evidence
FireRed-OCR
March 2026 release with repo, model card, and paper all live at launch
Stronger evidence package than many brand-new challengers, but still early for stability assumptions
DeepSeek-OCR-2
Strong HF traction soon after release
Good community momentum for tooling and examples
HunyuanOCR
High visibility and broad activity across channels
More examples in the wild for compact deployment patterns
GutenOCR
Growing technical interest from doc-AI builders
Strong relevance for grounding-heavy extraction workflows
5) A practical evaluation protocol (6 core models + challengers)
If you want one rigorous, reproducible process, run one fixed 50-page bake-off across six core models:
GutenOCR
HunyuanOCR
LightOnOCR-2-1B
DeepSeek-OCR-2
GLM-OCR
PaddleOCR-VL-1.5
Then add challenger tracks for newly released models. For this cycle:
dots.ocr-1.5 (especially if you need OCR plus web/screen/scene/SVG parsing)
FireRed-OCR (especially if malformed Markdown, broken formulas, or table closure failures are expensive in your workflow)
5.1 Preflight gates (before benchmarking)
Filter models before inference:
License/commercial gate
Region/compliance gate
Serving/runtime gate
Output-format gate
Practical note:
GutenOCR weights are CC-BY-NC, which often disqualifies direct commercial deployment.
HunyuanOCR uses a custom community license with territory and usage constraints, so legal review should happen before production rollout.
5.2 50-page stratified set
Slice
Pages
Why this slice matters
Clean digital single-column PDFs
8
Baseline text fidelity
Multi-column + sidebars + footnotes
8
Reading-order stress
Table-heavy documents
8
Structure fidelity and cell ordering
Formula-heavy documents
6
Formula extraction and sequencing
Forms/invoices/receipts
6
Region association and key-value linking
Messy photos/scans
10
Skew, warping, glare, and capture artifacts
Multilingual mixed-script pages
4
Language/layout stability
Total: 50 pages.
5.3 Ground-truth package per page
Prepare three artifacts for each page:
gt_text.txt
gt_markdown.md
gt_blocks.json with block_id, text, bbox, reading_index, and type
Quality control:
Dual-annotate all messy-photo pages and all multi-column pages.
Resolve disagreements before scoring.
5.4 Inference protocol (same policy for all models)
Render all pages at one fixed resolution (for example, 200 DPI).
Use deterministic decoding (temperature=0, no retries in the primary run).
Freeze model versions/commits and prompts.
Run one no-heuristic primary pass; report heuristic retries separately if used.
Mode recommendations:
DeepSeek-OCR-2: run both Free OCR and markdown conversion mode.
GLM-OCR: run markdown plus JSON layout output.
PaddleOCR-VL-1.5: run full document parsing mode.
dots.ocr-1.5: run document parsing mode first; if relevant, add web parsing and scene spotting prompts, and evaluate SVG output in a separate track.
FireRed-OCR: run its standard structured Markdown mode and score syntax-validity failures separately from plain text errors.
GutenOCR: run grounded mode (bbox outputs) and plain text mode.
HunyuanOCR: run document parsing prompt and spotting-style prompt where applicable.
LightOnOCR-2-1B: run standard OCR parsing mode.
5.5 Metrics
If this is your first pass, read this section as a checklist for formal evaluation teams. The important idea is simple: score text accuracy, layout order, table structure, and operational reliability separately so one strong number cannot hide a weak production behavior.
Reading-order metrics:
RO-ED (normalized reading-order edit distance, lower is better)
Kendall tau on reading_index sequence (higher is better)
Missing-block rate
Duplicate-block rate
Content and structure metrics:
CER and WER
Table TEDS
Formula metric (CDM or token-level F1, fixed across all runs)
Benchmark overfitting risk: do not promote a model to primary production without document-type stratified tests.
Layout drift risk: table structure quality can degrade faster than plain text quality across new templates.
Grounding risk: extraction pipelines fail when text is correct but linked to the wrong box or wrong row.
License risk: confirm commercial terms for each model/repo combination, not just the model card headline.
Operations risk: define fallback modes (text-only, markdown, or dual-model checks) before first rollout.
7) Conclusion
By February 2026, the market is no longer about finding one giant model to do everything. It is about matching the model to the failure mode you can least afford.
A practical rollout is:
Start with the use-case matrix in Section 1.
Shortlist three models with different strengths.
Run the fixed 50-page protocol.
Promote one primary model and one fallback model, then keep one fast-moving release lane for models like dots.ocr-1.5 or FireRed-OCR.
That is usually safer than picking one leaderboard winner and hoping the same order will hold on your own document mix.
Official March 2026 release frames it as the strongest end-to-end solution in its comparison slice; structural Markdown focus is the main differentiator