Recommendation for OCR & documents
OCR & Documents
Our top recommendation for OCR & Documents, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1][2] Use for high-stakes document OCR where human-rated image understanding quality matters, as it ranks in the top 5% of vision models on LMArena. DeepSeek: DeepSeek V4 Flash Vision Exp is the next-ranked alternative. Use via DeepSeek Harness or community plugins for free image transcription workflows with automatic vision failover to GLM-4V-Flash or Gemini.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 7
- Revision
- v77
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
38%
intended feed weight
Largest provider share
2 of 5
deepseek
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| VLMEvalKit tasksunavailable | 40% | feed unavailable | 0/20 |
| LMArena Document | 20% | #3 | 14/20 |
| LMArena Vision | 15% | #3 | 17/20 |
| Structured-output evalunavailable | 15% | feed unavailable | 0/20 |
| OpenRouter usage | 10% | 83/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- deepseek2 models
- Anthropic1 model
- Google1 model
- OpenAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Opus 4.6Anthropic | 54 | 38% | no linked practitioner threads | #3 LMArena Document · #3 LMArena Vision |
| 02 | DeepSeek V4 Flash Vision Expdeepseek | 47 | 11% | 2 threads · 2 families · 0 cautions | OpenRouter usage 86/100 normalized |
| 03 | DeepSeek V4 Flash 0423deepseek | 47 | 11% | 2 threads · 2 families · 1 cautions | OpenRouter usage 98/100 normalized |
| 04 | Gemini 3.5 Flash LiteGoogle | 47 | 28% | 1 threads · 1 families · 1 cautions | #20 LMArena Vision |
| 05 | GPT-5.6 LunaOpenAI | 46 | 38% | no linked practitioner threads | #16 LMArena Document · #26 LMArena Vision |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Opus 4.6 places third in human preference rankings for image understanding tasks, indicating strong visual reasoning capabilities relevant to document OCR.
Best when: Use for high-stakes document OCR where human-rated image understanding quality matters, as it ranks in the top 5% of vision models on LMArena.
Tips
- Use for high-stakes document OCR where human-rated image understanding quality matters, as it ranks in the top 5% of vision models on LMArena.
DeepSeek V4 Flash Vision Exp is an open-weight vision model with community tooling for image transcription, though evidence notes coordinate offset bugs in screen region selection.
Best when: Use via DeepSeek Harness or community plugins for free image transcription workflows with automatic vision failover to GLM-4V-Flash or Gemini.
Tips
- Use via DeepSeek Harness or community plugins for free image transcription workflows with automatic vision failover to GLM-4V-Flash or Gemini.
Watch out for
- Account for coordinate drift when selecting screen regions for OCR at 2560x1440 resolution with HDR enabled, as the actual capture area shifts left and up from the selection box.
DeepSeek V4 Flash 0423 shows specific failure modes on OCR-digitized government documents, with numeric extraction errors dominating its error taxonomy.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Avoid for Treasury Bulletin or similar financial document OCR, where 38 of 38 failures in one audit were numeric extraction and analysis errors after successful tool calls.
- Do not use the text-only 2.4T variant for OCR workflows, as the hosted API lacks vision capabilities and will fail silently on image inputs.
Gemini 3.5 Flash Lite ranks 23rd in vision preference with documented reliability issues on PDF OCR and frequent content rejection in automated document pipelines.
Best when: Switch to image mode rather than whole_pdf mode to reduce rejection rates when OCRing documents that trigger Gemini's content filters.
Tips
- Switch to image mode rather than whole_pdf mode to reduce rejection rates when OCRing documents that trigger Gemini's content filters.
Watch out for
- Expect frequent empty worker responses and API candidates with no content parts in paperless-gpt pipelines, requiring retry logic and causing workflow stalls.
- Watch for secondary rejection of OCR text during classification passes, where Gemini accepts the initial OCR but blocks the downstream LLM processing.
GPT-5.6 Luna ranks 28th in vision arena human preference, placing it in the bottom half of tested models for image understanding quality.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Consider alternatives for document OCR where image understanding accuracy is critical, as its LMArena Elo of 1258 trails 27 vision models including several cheaper options.
Frequently asked
- What is the top-ranked model for OCR & Documents?
- Anthropic: Claude Opus 4.6 ranks first in the current evidence-weighted comparison. Use for high-stakes document OCR where human-rated image understanding quality matters, as it ranks in the top 5% of vision models on LMArena.[1]
- What is an alternative to Anthropic: Claude Opus 4.6?
- DeepSeek: DeepSeek V4 Flash Vision Exp is the next-ranked option. Use via DeepSeek Harness or community plugins for free image transcription workflows with automatic vision failover to GLM-4V-Flash or Gemini.[2]
Sources
- 1
“Ranks #3 of 71 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Sep 13, 2026 - 2
“### Kind plugin ### Name dsh-media-skills ### GitHub repository MJorgin/dsh-media-skills ### npm package _No response_ ### Description Free image reading & generation for DeepSeek Harness (rc.7 / rc.8 / v0.1.1-rc.1 / rc.2) — paste-image reading with auto vision transcription, DeepSeek-V4-Flash-Vision-Exp / GLM-4V-Flash / SenseNova / Gemini failover, Kolors + U1 Fast generation. No keys in repo. ### Tags dsh-plugin, community ### Submitter GitHub username dsh-registry-bot”
deepseek-harness-plugin[bot] · GitHub · Aug 24, 2026 - 3
“OCR选择屏幕区域时,实际的区域要比框选的区域偏左上一些,屏幕分辨率为2560x1440,开启HDR 因为Windows本地OCR识别率偏低,故使用deepseek-v4-flash-vision-exp模型测试”
DQ-HollyLee · GitHub · Sep 8, 2026 - 4
“# Daily ironclaw failure taxonomy — 2026-09-24 ## Suites analyzed - [officeqa (38 non-pass)](https://nearai.github.io/benchmarks/#/runs/ironclaw/officeqa/69f799a5-16f3-4577-b008-89c48b9deb82) — All 38 non-pass tasks are genuine model-quality failures by deepseek-v4-flash over OCR-digitized Treasury Bulletins; nothing in ironclaw fails. The dominant mode (recurring from the prior run) is numeric extraction/analysis error: the agent's read/grep/shell/python tool calls all succeed and return the c…”
pranavraja99 · GitHub · Sep 24, 2026 - 5
“## Standing order Do **not** OCR-then-DeepSeek / text-only 2.4T. DeepSeek v4-flash was unscored because the hosted API is text-only. If a VL host for Qwen3.8-2.4T-A95B appears, run the same seed-777 JSONL + scorer as Luna/restem. Until then this is parked.”
caiotheodoro · GitHub · Aug 21, 2026 - 6
“1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…”
phaset · GitHub · Jul 28, 2026 - 7
“Ranks #28 of 71 on LMArena's vision arena (Elo 1258), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Sep 13, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.