Recommendation for OCR & documents

OCR & Documents

Our top recommendation for OCR & Documents, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1][2] Use for high-stakes document OCR where human-rated image understanding quality matters, as it ranks in the top 5% of vision models on LMArena. DeepSeek: DeepSeek V4 Flash Vision Exp is the next-ranked alternative. Use via DeepSeek Harness or community plugins for free image transcription workflows with automatic vision failover to GLM-4V-Flash or Gemini.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
7
Revision
v77

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

38%

intended feed weight

Largest provider share

2 of 5

deepseek

Provisional source breadth. 6 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 38%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Opus 4.6
Evaluation feedWeightWinner resultField measured
VLMEvalKit tasksunavailable
40%
feed unavailable0/20
LMArena Document
20%
#314/20
LMArena Vision
15%
#317/20
Structured-output evalunavailable
15%
feed unavailable0/20
OpenRouter usage
10%
83/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

deepseek40%
  • deepseek2 models
  • Anthropic1 model
  • Google1 model
  • OpenAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Opus 4.6Anthropic
54
38%no linked practitioner threads#3 LMArena Document · #3 LMArena Vision
02DeepSeek V4 Flash Vision Expdeepseek
47
11%2 threads · 2 families · 0 cautionsOpenRouter usage 86/100 normalized
03DeepSeek V4 Flash 0423deepseek
47
11%2 threads · 2 families · 1 cautionsOpenRouter usage 98/100 normalized
04Gemini 3.5 Flash LiteGoogle
47
28%1 threads · 1 families · 1 cautions#20 LMArena Vision
05GPT-5.6 LunaOpenAI
46
38%no linked practitioner threads#16 LMArena Document · #26 LMArena Vision

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Opus 4.6 places third in human preference rankings for image understanding tasks, indicating strong visual reasoning capabilities relevant to document OCR.

    Best when: Use for high-stakes document OCR where human-rated image understanding quality matters, as it ranks in the top 5% of vision models on LMArena.

    Tips

    • Use for high-stakes document OCR where human-rated image understanding quality matters, as it ranks in the top 5% of vision models on LMArena.
      Source 1
      “Ranks #3 of 71 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.”
      LMArena vision arenaOpen original ↗
  2. DeepSeek V4 Flash Vision Exp is an open-weight vision model with community tooling for image transcription, though evidence notes coordinate offset bugs in screen region selection.

    Best when: Use via DeepSeek Harness or community plugins for free image transcription workflows with automatic vision failover to GLM-4V-Flash or Gemini.

    Tips

    • Use via DeepSeek Harness or community plugins for free image transcription workflows with automatic vision failover to GLM-4V-Flash or Gemini.
      Source 2
      “### Kind plugin ### Name dsh-media-skills ### GitHub repository MJorgin/dsh-media-skills ### npm package _No response_ ### Description Free image reading & generation for DeepSeek Harness (rc.7 / rc.8 / v0.1.1-rc.1 / rc.2) — paste-image reading with auto vision transcription, DeepSeek-V4-Flash-Vision-Exp / GLM-4V-Flash / SenseNova / Gemini failover, Kolors + U1 Fast generation. No keys in repo. ### Tags dsh-plugin, community ### Submitter GitHub username dsh-registry-bot”
      deepseek-harness-plugin[bot]Open original ↗

    Watch out for

    • Account for coordinate drift when selecting screen regions for OCR at 2560x1440 resolution with HDR enabled, as the actual capture area shifts left and up from the selection box.
      Source 3
      “OCR选择屏幕区域时,实际的区域要比框选的区域偏左上一些,屏幕分辨率为2560x1440,开启HDR 因为Windows本地OCR识别率偏低,故使用deepseek-v4-flash-vision-exp模型测试”
  3. DeepSeek V4 Flash 0423 shows specific failure modes on OCR-digitized government documents, with numeric extraction errors dominating its error taxonomy.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Avoid for Treasury Bulletin or similar financial document OCR, where 38 of 38 failures in one audit were numeric extraction and analysis errors after successful tool calls.
      Source 4
      “# Daily ironclaw failure taxonomy — 2026-09-24 ## Suites analyzed - [officeqa (38 non-pass)](https://nearai.github.io/benchmarks/#/runs/ironclaw/officeqa/69f799a5-16f3-4577-b008-89c48b9deb82) — All 38 non-pass tasks are genuine model-quality failures by deepseek-v4-flash over OCR-digitized Treasury Bulletins; nothing in ironclaw fails. The dominant mode (recurring from the prior run) is numeric extraction/analysis error: the agent's read/grep/shell/python tool calls all succeed and return the c…”
      pranavraja99Open original ↗
    • Do not use the text-only 2.4T variant for OCR workflows, as the hosted API lacks vision capabilities and will fail silently on image inputs.
      Source 5
      “## Standing order Do **not** OCR-then-DeepSeek / text-only 2.4T. DeepSeek v4-flash was unscored because the hosted API is text-only. If a VL host for Qwen3.8-2.4T-A95B appears, run the same seed-777 JSONL + scorer as Luna/restem. Until then this is parked.”
      caiotheodoroOpen original ↗
  4. Gemini 3.5 Flash Lite ranks 23rd in vision preference with documented reliability issues on PDF OCR and frequent content rejection in automated document pipelines.

    Best when: Switch to image mode rather than whole_pdf mode to reduce rejection rates when OCRing documents that trigger Gemini's content filters.

    Tips

    • Switch to image mode rather than whole_pdf mode to reduce rejection rates when OCRing documents that trigger Gemini's content filters.
      Source 6
      “1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…”

    Watch out for

    • Expect frequent empty worker responses and API candidates with no content parts in paperless-gpt pipelines, requiring retry logic and causing workflow stalls.
      Source 6
      “1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…”
    • Watch for secondary rejection of OCR text during classification passes, where Gemini accepts the initial OCR but blocks the downstream LLM processing.
      Source 6
      “1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…”
  5. GPT-5.6 Luna ranks 28th in vision arena human preference, placing it in the bottom half of tested models for image understanding quality.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Consider alternatives for document OCR where image understanding accuracy is critical, as its LMArena Elo of 1258 trails 27 vision models including several cheaper options.
      Source 7
      “Ranks #28 of 71 on LMArena's vision arena (Elo 1258), based on human preference on image-understanding tasks.”
      LMArena vision arenaOpen original ↗

Frequently asked

What is the top-ranked model for OCR & Documents?
Anthropic: Claude Opus 4.6 ranks first in the current evidence-weighted comparison. Use for high-stakes document OCR where human-rated image understanding quality matters, as it ranks in the top 5% of vision models on LMArena.[1]
What is an alternative to Anthropic: Claude Opus 4.6?
DeepSeek: DeepSeek V4 Flash Vision Exp is the next-ranked option. Use via DeepSeek Harness or community plugins for free image transcription workflows with automatic vision failover to GLM-4V-Flash or Gemini.[2]

Sources

  1. 1

    “Ranks #3 of 71 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.”

    LMArena vision arena · Benchmark · Sep 13, 2026
  2. 2

    “### Kind plugin ### Name dsh-media-skills ### GitHub repository MJorgin/dsh-media-skills ### npm package _No response_ ### Description Free image reading & generation for DeepSeek Harness (rc.7 / rc.8 / v0.1.1-rc.1 / rc.2) — paste-image reading with auto vision transcription, DeepSeek-V4-Flash-Vision-Exp / GLM-4V-Flash / SenseNova / Gemini failover, Kolors + U1 Fast generation. No keys in repo. ### Tags dsh-plugin, community ### Submitter GitHub username dsh-registry-bot”

    deepseek-harness-plugin[bot] · GitHub · Aug 24, 2026
  3. 3

    “OCR选择屏幕区域时,实际的区域要比框选的区域偏左上一些,屏幕分辨率为2560x1440,开启HDR 因为Windows本地OCR识别率偏低,故使用deepseek-v4-flash-vision-exp模型测试”

    DQ-HollyLee · GitHub · Sep 8, 2026
  4. 4

    “# Daily ironclaw failure taxonomy — 2026-09-24 ## Suites analyzed - [officeqa (38 non-pass)](https://nearai.github.io/benchmarks/#/runs/ironclaw/officeqa/69f799a5-16f3-4577-b008-89c48b9deb82) — All 38 non-pass tasks are genuine model-quality failures by deepseek-v4-flash over OCR-digitized Treasury Bulletins; nothing in ironclaw fails. The dominant mode (recurring from the prior run) is numeric extraction/analysis error: the agent's read/grep/shell/python tool calls all succeed and return the c…”

    pranavraja99 · GitHub · Sep 24, 2026
  5. 5

    “## Standing order Do **not** OCR-then-DeepSeek / text-only 2.4T. DeepSeek v4-flash was unscored because the hosted API is text-only. If a VL host for Qwen3.8-2.4T-A95B appears, run the same seed-777 JSONL + scorer as Luna/restem. Until then this is parked.”

    caiotheodoro · GitHub · Aug 21, 2026
  6. 6

    “1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…”

    phaset · GitHub · Jul 28, 2026
  7. 7

    “Ranks #28 of 71 on LMArena's vision arena (Elo 1258), based on human preference on image-understanding tasks.”

    LMArena vision arena · Benchmark · Sep 13, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.