Which pages favour Hunyuan over FireRed, and vice versa?

By Wei Jie Chee · 28 Mar 2026, 00:00 Z

Download printable cheat-sheet (CC-BY 4.0)
Named-model investigator + Constrained implementerRead bounded results for individual model choices.

Hunyuan's grounding helps on some faint, formula, and blank pages, while FireRed's Markdown-first path is faster and stronger on several structured page types.

Do this first

Separate faint, formula, blank, and structured pages, then compare quality and runtime by page type instead of averaging them together.

Keep this boundary in view

Cross-model consensus is a practical reference, not manually transcribed ground truth, and neither model wins every page type.

Inspect the evidence and what would change the conclusion
Evidence layer
Hunyuan and FireRed full-50 comparison
Observed
March 2026
Sample
50 pages across seven types; 49 processed per workflow

What supports it

The comparison used the same 50-page, seven-type slice; each workflow processed 49 pages.

Change our mind if

A manually adjudicated held-out set removes the observed page-type split or reverses the operational trade-off.

Claim OCR-HUNYUAN-FIRERED-50-V1 / v1, checked 6 Aug 2026

This post answers a head-to-head question: if you are choosing between Hunyuan OCR and FireRed OCR, which one should handle which scanned pages.

The short version

  • Neither Hunyuan OCR nor FireRed OCR is universally better.
  • Across 50 scanned pages in 7 page types, Hunyuan won 4 page types and FireRed won 3.
  • Hunyuan is stronger when the workflow needs coordinate-grounded output or messy-page reasoning.
  • FireRed is faster and cleaner when the page already has clear structure.
  • The practical answer is to route by page type, not pick one model globally.

The one-minute decision path

The split is architectural. Hunyuan returns text with page coordinates, so it can help on degraded, ambiguous, or blank pages where spatial evidence matters. FireRed produces markdown-first output, so it is faster and cleaner on structured content that already reads clearly.

If the page is...Start with...Why
low-contrast, faint, blank, or formula-heavyHunyuan OCRcoordinate-grounded output gives better evidence on harder pages
table-heavy, worksheet-style, or diagram-questionFireRed OCRmarkdown-first output was cleaner or faster in these slices
text-first noteseither, then optimize for speed or cleanup costthe measured gap was small

Local test / OCR-HUNYUAN-FIRERED-50-V1

Test this claim on your documents.

Copy a read-only experiment brief that asks your coding agent to preserve raw evidence, test a negative case, and return a verdict that can disagree with this page.

Follow OCR research via RSS

Preview the experiment brief
Use Codex to test this bounded OCR claim in my local environment.

Source: https://instavar.com/research/ocr/hunyuan-ocr-vs-firered-ocr-2026
Claim ID: OCR-HUNYUAN-FIRERED-50-V1
Claim version: v1, checked 6 Aug 2026
Claim: Hunyuan's grounding helps on some faint, formula, and blank pages, while FireRed's Markdown-first path is faster and stronger on several structured page types.
Observed evidence: The comparison used the same 50-page, seven-type slice; each workflow processed 49 pages.
Evidence layer: Hunyuan and FireRed full-50 comparison.
Sample: 50 pages across seven types; 49 processed per workflow.
Boundary: Cross-model consensus is a practical reference, not manually transcribed ground truth, and neither model wins every page type.
Potential falsifier: A manually adjudicated held-out set removes the observed page-type split or reverses the operational trade-off.

Safety and privacy:
- Work read-only by default and keep all documents and outputs local.
- Do not upload documents, reveal credentials, or send document contents to an external service.
- Do not install software, download large models, start paid jobs, or change my environment without asking first.
- If the available tools cannot run the comparison safely, stop with a concrete setup plan instead of pretending the claim was tested.

Test method:
1. Inspect the available OCR workflows and record their exact versions, runtime, hardware, and configuration.
2. Help me choose a small representative sample. Include a known difficult page and a negative control with expected text or structure when available.
3. Before running anything, state the primary failure we are testing and the measurement that would expose it.
4. Preserve raw outputs. Compare text fidelity, structure, blank-page behavior, grounding correctness, runtime, and cleanup effort only where those properties can actually be measured.
5. Treat anchor count as an unverified quantity until sampled anchors are checked against the page.
6. Report the result as support, contradict, mixed, or inconclusive. Explain what observation produced that verdict.
7. Name the smallest follow-up test that could change the verdict.
What did your local test say?

Save a coarse verdict in this browser. Instavar receives no document contents through this control.