Recommendation for Documents & forms

Document Parsing

Our top recommendation for Document Parsing, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1] Anthropic: Claude Fable 5 is the next-ranked alternative. Deploy for general document parsing workflows where its 1496 document arena score indicates strong blind preference performance.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
10
Revision
v79

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

38%

intended feed weight

Largest provider share

2 of 6

Anthropic

Provisional source breadth. 7 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 33%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Opus 4.6
Evaluation feedWeightWinner resultField measured
VLMEvalKit tasksunavailable
40%
feed unavailable0/20
LMArena Document
20%
#314/20
LMArena Vision
15%
#318/20
Structured-output evalunavailable
15%
feed unavailable0/20
OpenRouter usage
10%
83/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic33%
  • Anthropic2 models
  • OpenAI2 models
  • deepseek1 model
  • Google1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Opus 4.6Anthropic
54
38%no linked practitioner threads#3 LMArena Document · #3 LMArena Vision
02Claude Fable 5Anthropic
53
38%1 threads · 1 families · 1 cautions#1 LMArena Vision · #4 LMArena Document
03GPT-5.6 SolOpenAI
52
38%2 threads · 2 families · 1 cautions#6 LMArena Document · #8 LMArena Vision
04GPT-5.6 LunaOpenAI
50
38%2 threads · 2 families · 0 cautions#16 LMArena Document · #26 LMArena Vision
05DeepSeek V4 Flash Vision Expdeepseek
47
11%4 threads · 3 families · 1 cautionsOpenRouter usage 86/100 normalized
06Gemini 2.5 FlashGoogle
47
28%2 threads · 2 families · 0 cautions#40 LMArena Vision

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Anthropic: Claude Opus 4.6 ranks #3 of 34 on LMArena's document arena (score 1507), based on blind preference for document tasks.

    Best when: Consider only after reviewing the cited caution.

  2. Claude Fable 5 ranks fourth on LMArena's document arena with a 1496 score, though evidence notes it refuses the majority of questions in specialized evaluations like LifeSciBench and MedChemBench.

    Best when: Deploy for general document parsing workflows where its 1496 document arena score indicates strong blind preference performance.

    Tips

    • Deploy for general document parsing workflows where its 1496 document arena score indicates strong blind preference performance.
      Source 1
      “Ranks #4 of 34 on LMArena's document arena (score 1496), based on blind preference for document tasks.”
      LMArena document arenaOpen original ↗

    Watch out for

    • Expect frequent refusals on domain-specific content; the model was excluded from multiple scientific benchmarks because it refuses most questions.
      Source 2
      “"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12" Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.”
  3. GPT-5.6 Sol ranks ninth on LMArena's vision arena with an Elo of 1286, with community evidence showing it handles hex decoding from packet dumps and decompilation tasks with tool assistance.

    Best when: Deploy for technical document parsing involving encoded or binary data, as it has handled hex decoding from packet dumps and buffer offset searches in security contexts.

    Tips

    • Deploy for technical document parsing involving encoded or binary data, as it has handled hex decoding from packet dumps and buffer offset searches in security contexts.
      Source 3
      “It's good enough at decoding hex from some packet dumps. And I was doing that even with 5.5. And it was good at decompiling some code (with tools) and searching for offsets of buffers and commands. Found viable exploit that allowed me to rescue broken update system in devices I was maintaining for my company (it was broken by chatgpt forgetting -v in hexdump, heh).”
  4. GPT-5.6 Luna ranks 28th on LMArena's vision arena with an Elo of 1258, with production evidence positioning it as a $4 fallback option for document backlog processing when primary models hit rate limits.

    Best when: Keep as a fallback for document extraction pipelines, with production evidence showing it costs approximately $4 for full backlog processing when primary models are unavailable.

    Tips

    • Keep as a fallback for document extraction pipelines, with production evidence showing it costs approximately $4 for full backlog processing when primary models are unavailable.
      Source 4
      “> *This was generated by AI during triage.* ## Parent #205 ## What to build A client module that takes document text and returns validated extraction records conforming to the schema in `docs/protocols/extraction-schema-by-document-type.md`. Two models: **Nemotron 3 Ultra (free)** as default, **GPT-5.6-Luna (~$4 for the full backlog)** as automatic fallback on rate-limit or unavailability. Via OpenRouter. Four things that already bit during evaluation and must be handled: 1. **Retry with backof…”

    Watch out for

    • Expect lower vision quality than alternatives; its 1258 Elo places it in the bottom half of LMArena's vision arena rankings.
      Source 5
      “Ranks #28 of 71 on LMArena's vision arena (Elo 1258), based on human preference on image-understanding tasks.”
      LMArena vision arenaOpen original ↗
  5. DeepSeek V4 Flash Vision Exp is an open-weight model with production evidence showing 9–13 second latency per successful visual qualification call and 95 seconds per presentation generation call, though community notes suggest it is not designed for lossless table extraction from screenshots.

    Best when: Use for multi-page document pipelines requiring strict schema extraction with observable timeout/retry/fallback patterns, as deployed in production specs.

    Tips

    • Use for multi-page document pipelines requiring strict schema extraction with observable timeout/retry/fallback patterns, as deployed in production specs.
      Source 6
      “## Objetivo Reducir el tiempo real de calificación, digitalización y generación de presentaciones sin sacrificar calidad, trazabilidad ni perder solicitudes en curso. ## Evidencia de producción (últimos 14 días) - Calificación completa: p50 157 s; p95 516 s. - Digitalización: p50 189 s. - Presentación reciente: 366 s. - DeepSeek V4 Flash Vision Exp en calificación visual: ~9–13 s por llamada exitosa. - Presentaciones con ese modelo: ~95 s por llamada y hasta tres llamadas por regeneración/revis…”
      Source 7
      “Implementar specs/020-deepseek-vision: extractor desacoplado, multipágina, schema estricto, timeout/retry/fallback observable y benchmarks directos/backend. La solicitud detallada del usuario constituye aprobación de spec y plan.”
    • Select when you need an open-weight vision model with measured latency: ~9–13 seconds per visual qualification call in production workloads.
      Source 6
      “## Objetivo Reducir el tiempo real de calificación, digitalización y generación de presentaciones sin sacrificar calidad, trazabilidad ni perder solicitudes en curso. ## Evidencia de producción (últimos 14 días) - Calificación completa: p50 157 s; p95 516 s. - Digitalización: p50 189 s. - Presentación reciente: 366 s. - DeepSeek V4 Flash Vision Exp en calificación visual: ~9–13 s por llamada exitosa. - Presentaciones con ese modelo: ~95 s por llamada y hasta tres llamadas por regeneración/revis…”

    Watch out for

    • Do not rely on for fine-grained lossless table extraction from screenshots; community evidence indicates generalist vision models are not built for resolving grid cells back to tables.
      Source 8
      “I think there's an expectation mismatch here. If you're trying to do fine-grained data extraction off a screenshot, you'd be better off trying an OCR model. Generalist vision-models aren't made for losslessly resolving grid cells back to a table.”
  6. Gemini 2.5 Flash is documented for structured JSON output from release notes parsing, though evidence indicates it faces scheduled deprecation with ELA entry in October 2026 and potential price increases in January 2027.

    Best when: Use for schema-bound extraction tasks where structured output and JSON Schema compliance are required, as shown in release note decomposition workflows.

    Tips

    • Use for schema-bound extraction tasks where structured output and JSON Schema compliance are required, as shown in release note decomposition workflows.
      Source 9
      “## 概要 GitHub Releaseの原文をGemini APIで解析し、変更単位の日本語要約・Breaking Change・Impact・Migrationを構造化して保存する。 ## 対応内容 - [ ] Gemini API連携 - [ ] `gemini-2.5-flash` 利用 - [ ] JSON Schema / structured output - [ ] Release Notesを変更単位へ分解 - [ ] 日本語タイトル・要約 - [ ] category判定 - [ ] impact判定 - [ ] breaking判定 - [ ] migration判定 - [ ] Gemini API失敗時のリトライ - [ ] APIキーをGitHub Secretsで管理 - [ ] Phase 2の自動収集との統合 - [ ] テスト ## 出力 `changes`, `category`, `title`, `summary`, `impact`, `breaking`, `migration` を構造化JSONとして保存する。 ## 完了条件 - Re…”

    Watch out for

    • Plan migration before October 2026; the model enters ELA with continued availability but faces significant price changes and regional availability shifts after January 2027.
      Source 10
      “## 概要 Vertex AI(Gemini Enterprise Agent Platform)の `gemini-2.5-flash` が 2026-10-20 に廃止(ELA 入り)予定のため、Gemini 3 系への移行先モデルを実測で選定し、ADR にまとめる。 Google からの通知では影響プロジェクトとして `documentaisample-488504` が名指しされており、本リポジトリの Gemini 抽出経路が対象。 - 2026-10-20: ELA 入り(呼び出しは継続可能・価格据え置き) - 2027-01-28: ELA 標準価格終了、大幅値上げ+リージョン提供状況の変更可能性 呼び出しが即停止するわけではないため障害リスクは低いが、ADR-0010 が記録したコスト前提(`thinkingBudget:0` での平均 $0.00175/枚)が失効するため、実測をやり直す必要がある。 ## 変更内容 本 Issue のスコープは **移行先モデルの実測比較と方針決定(ADR)まで**。既定モデルの切替コードは別 Issue で対応する。 1. 移行先候…”

Frequently asked

What is an alternative to Anthropic: Claude Opus 4.6?
Anthropic: Claude Fable 5 is the next-ranked option. Deploy for general document parsing workflows where its 1496 document arena score indicates strong blind preference performance.[1]

Sources

  1. 1

    “Ranks #4 of 34 on LMArena's document arena (score 1496), based on blind preference for document tasks.”

    LMArena document arena · Benchmark · Sep 13, 2026
  2. 2

    “"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12" Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.”

    HDBaseT · Hacker News · Sep 3, 2026
  3. 3

    “It's good enough at decoding hex from some packet dumps. And I was doing that even with 5.5. And it was good at decompiling some code (with tools) and searching for offsets of buffers and commands. Found viable exploit that allowed me to rescue broken update system in devices I was maintaining for my company (it was broken by chatgpt forgetting -v in hexdump, heh).”

    yetihehe · Hacker News · Aug 13, 2026
  4. 4

    “> *This was generated by AI during triage.* ## Parent #205 ## What to build A client module that takes document text and returns validated extraction records conforming to the schema in `docs/protocols/extraction-schema-by-document-type.md`. Two models: **Nemotron 3 Ultra (free)** as default, **GPT-5.6-Luna (~$4 for the full backlog)** as automatic fallback on rate-limit or unavailability. Via OpenRouter. Four things that already bit during evaluation and must be handled: 1. **Retry with backof…”

    alexwolson · GitHub · Aug 6, 2026
  5. 5

    “Ranks #28 of 71 on LMArena's vision arena (Elo 1258), based on human preference on image-understanding tasks.”

    LMArena vision arena · Benchmark · Sep 13, 2026
  6. 6

    “## Objetivo Reducir el tiempo real de calificación, digitalización y generación de presentaciones sin sacrificar calidad, trazabilidad ni perder solicitudes en curso. ## Evidencia de producción (últimos 14 días) - Calificación completa: p50 157 s; p95 516 s. - Digitalización: p50 189 s. - Presentación reciente: 366 s. - DeepSeek V4 Flash Vision Exp en calificación visual: ~9–13 s por llamada exitosa. - Presentaciones con ese modelo: ~95 s por llamada y hasta tres llamadas por regeneración/revis…”

    Andres-back · GitHub · Sep 4, 2026
  7. 7

    “Implementar specs/020-deepseek-vision: extractor desacoplado, multipágina, schema estricto, timeout/retry/fallback observable y benchmarks directos/backend. La solicitud detallada del usuario constituye aprobación de spec y plan.”

    Andres-back · GitHub · Aug 24, 2026
  8. 8

    “I think there's an expectation mismatch here. If you're trying to do fine-grained data extraction off a screenshot, you'd be better off trying an OCR model. Generalist vision-models aren't made for losslessly resolving grid cells back to a table.”

    Skelectric · GitHub · Sep 17, 2026
  9. 9

    “## 概要 GitHub Releaseの原文をGemini APIで解析し、変更単位の日本語要約・Breaking Change・Impact・Migrationを構造化して保存する。 ## 対応内容 - [ ] Gemini API連携 - [ ] `gemini-2.5-flash` 利用 - [ ] JSON Schema / structured output - [ ] Release Notesを変更単位へ分解 - [ ] 日本語タイトル・要約 - [ ] category判定 - [ ] impact判定 - [ ] breaking判定 - [ ] migration判定 - [ ] Gemini API失敗時のリトライ - [ ] APIキーをGitHub Secretsで管理 - [ ] Phase 2の自動収集との統合 - [ ] テスト ## 出力 `changes`, `category`, `title`, `summary`, `impact`, `breaking`, `migration` を構造化JSONとして保存する。 ## 完了条件 - Re…”

    haruki33 · GitHub · Aug 17, 2026
  10. 10

    “## 概要 Vertex AI(Gemini Enterprise Agent Platform)の `gemini-2.5-flash` が 2026-10-20 に廃止(ELA 入り)予定のため、Gemini 3 系への移行先モデルを実測で選定し、ADR にまとめる。 Google からの通知では影響プロジェクトとして `documentaisample-488504` が名指しされており、本リポジトリの Gemini 抽出経路が対象。 - 2026-10-20: ELA 入り(呼び出しは継続可能・価格据え置き) - 2027-01-28: ELA 標準価格終了、大幅値上げ+リージョン提供状況の変更可能性 呼び出しが即停止するわけではないため障害リスクは低いが、ADR-0010 が記録したコスト前提(`thinkingBudget:0` での平均 $0.00175/枚)が失効するため、実測をやり直す必要がある。 ## 変更内容 本 Issue のスコープは **移行先モデルの実測比較と方針決定(ADR)まで**。既定モデルの切替コードは別 Issue で対応する。 1. 移行先候…”

    git-berian · GitHub · Aug 7, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.