Recommendation for Documents & forms
Document Parsing
Our top recommendation for Document Parsing, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1] Anthropic: Claude Fable 5 is the next-ranked alternative. Deploy for general document parsing workflows where its 1496 document arena score indicates strong blind preference performance.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 10
- Revision
- v79
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
38%
intended feed weight
Largest provider share
2 of 6
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| VLMEvalKit tasksunavailable | 40% | feed unavailable | 0/20 |
| LMArena Document | 20% | #3 | 14/20 |
| LMArena Vision | 15% | #3 | 18/20 |
| Structured-output evalunavailable | 15% | feed unavailable | 0/20 |
| OpenRouter usage | 10% | 83/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic2 models
- OpenAI2 models
- deepseek1 model
- Google1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Opus 4.6Anthropic | 54 | 38% | no linked practitioner threads | #3 LMArena Document · #3 LMArena Vision |
| 02 | Claude Fable 5Anthropic | 53 | 38% | 1 threads · 1 families · 1 cautions | #1 LMArena Vision · #4 LMArena Document |
| 03 | GPT-5.6 SolOpenAI | 52 | 38% | 2 threads · 2 families · 1 cautions | #6 LMArena Document · #8 LMArena Vision |
| 04 | GPT-5.6 LunaOpenAI | 50 | 38% | 2 threads · 2 families · 0 cautions | #16 LMArena Document · #26 LMArena Vision |
| 05 | DeepSeek V4 Flash Vision Expdeepseek | 47 | 11% | 4 threads · 3 families · 1 cautions | OpenRouter usage 86/100 normalized |
| 06 | Gemini 2.5 FlashGoogle | 47 | 28% | 2 threads · 2 families · 0 cautions | #40 LMArena Vision |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Anthropic: Claude Opus 4.6 ranks #3 of 34 on LMArena's document arena (score 1507), based on blind preference for document tasks.
Best when: Consider only after reviewing the cited caution.
Claude Fable 5 ranks fourth on LMArena's document arena with a 1496 score, though evidence notes it refuses the majority of questions in specialized evaluations like LifeSciBench and MedChemBench.
Best when: Deploy for general document parsing workflows where its 1496 document arena score indicates strong blind preference performance.
Tips
- Deploy for general document parsing workflows where its 1496 document arena score indicates strong blind preference performance.
Watch out for
- Expect frequent refusals on domain-specific content; the model was excluded from multiple scientific benchmarks because it refuses most questions.
GPT-5.6 Sol ranks ninth on LMArena's vision arena with an Elo of 1286, with community evidence showing it handles hex decoding from packet dumps and decompilation tasks with tool assistance.
Best when: Deploy for technical document parsing involving encoded or binary data, as it has handled hex decoding from packet dumps and buffer offset searches in security contexts.
Tips
- Deploy for technical document parsing involving encoded or binary data, as it has handled hex decoding from packet dumps and buffer offset searches in security contexts.
GPT-5.6 Luna ranks 28th on LMArena's vision arena with an Elo of 1258, with production evidence positioning it as a $4 fallback option for document backlog processing when primary models hit rate limits.
Best when: Keep as a fallback for document extraction pipelines, with production evidence showing it costs approximately $4 for full backlog processing when primary models are unavailable.
Tips
- Keep as a fallback for document extraction pipelines, with production evidence showing it costs approximately $4 for full backlog processing when primary models are unavailable.
Watch out for
- Expect lower vision quality than alternatives; its 1258 Elo places it in the bottom half of LMArena's vision arena rankings.
DeepSeek V4 Flash Vision Exp is an open-weight model with production evidence showing 9–13 second latency per successful visual qualification call and 95 seconds per presentation generation call, though community notes suggest it is not designed for lossless table extraction from screenshots.
Best when: Use for multi-page document pipelines requiring strict schema extraction with observable timeout/retry/fallback patterns, as deployed in production specs.
Tips
- Use for multi-page document pipelines requiring strict schema extraction with observable timeout/retry/fallback patterns, as deployed in production specs.
- Select when you need an open-weight vision model with measured latency: ~9–13 seconds per visual qualification call in production workloads.
Watch out for
- Do not rely on for fine-grained lossless table extraction from screenshots; community evidence indicates generalist vision models are not built for resolving grid cells back to tables.
Gemini 2.5 Flash is documented for structured JSON output from release notes parsing, though evidence indicates it faces scheduled deprecation with ELA entry in October 2026 and potential price increases in January 2027.
Best when: Use for schema-bound extraction tasks where structured output and JSON Schema compliance are required, as shown in release note decomposition workflows.
Tips
- Use for schema-bound extraction tasks where structured output and JSON Schema compliance are required, as shown in release note decomposition workflows.
Watch out for
- Plan migration before October 2026; the model enters ELA with continued availability but faces significant price changes and regional availability shifts after January 2027.
Frequently asked
- What is an alternative to Anthropic: Claude Opus 4.6?
- Anthropic: Claude Fable 5 is the next-ranked option. Deploy for general document parsing workflows where its 1496 document arena score indicates strong blind preference performance.[1]
Sources
- 1
“Ranks #4 of 34 on LMArena's document arena (score 1496), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Sep 13, 2026 - 2
“"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12" Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.”
HDBaseT · Hacker News · Sep 3, 2026 - 3
“It's good enough at decoding hex from some packet dumps. And I was doing that even with 5.5. And it was good at decompiling some code (with tools) and searching for offsets of buffers and commands. Found viable exploit that allowed me to rescue broken update system in devices I was maintaining for my company (it was broken by chatgpt forgetting -v in hexdump, heh).”
yetihehe · Hacker News · Aug 13, 2026 - 4
“> *This was generated by AI during triage.* ## Parent #205 ## What to build A client module that takes document text and returns validated extraction records conforming to the schema in `docs/protocols/extraction-schema-by-document-type.md`. Two models: **Nemotron 3 Ultra (free)** as default, **GPT-5.6-Luna (~$4 for the full backlog)** as automatic fallback on rate-limit or unavailability. Via OpenRouter. Four things that already bit during evaluation and must be handled: 1. **Retry with backof…”
alexwolson · GitHub · Aug 6, 2026 - 5
“Ranks #28 of 71 on LMArena's vision arena (Elo 1258), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Sep 13, 2026 - 6
“## Objetivo Reducir el tiempo real de calificación, digitalización y generación de presentaciones sin sacrificar calidad, trazabilidad ni perder solicitudes en curso. ## Evidencia de producción (últimos 14 días) - Calificación completa: p50 157 s; p95 516 s. - Digitalización: p50 189 s. - Presentación reciente: 366 s. - DeepSeek V4 Flash Vision Exp en calificación visual: ~9–13 s por llamada exitosa. - Presentaciones con ese modelo: ~95 s por llamada y hasta tres llamadas por regeneración/revis…”
Andres-back · GitHub · Sep 4, 2026 - 7
“Implementar specs/020-deepseek-vision: extractor desacoplado, multipágina, schema estricto, timeout/retry/fallback observable y benchmarks directos/backend. La solicitud detallada del usuario constituye aprobación de spec y plan.”
Andres-back · GitHub · Aug 24, 2026 - 8
“I think there's an expectation mismatch here. If you're trying to do fine-grained data extraction off a screenshot, you'd be better off trying an OCR model. Generalist vision-models aren't made for losslessly resolving grid cells back to a table.”
Skelectric · GitHub · Sep 17, 2026 - 9
“## 概要 GitHub Releaseの原文をGemini APIで解析し、変更単位の日本語要約・Breaking Change・Impact・Migrationを構造化して保存する。 ## 対応内容 - [ ] Gemini API連携 - [ ] `gemini-2.5-flash` 利用 - [ ] JSON Schema / structured output - [ ] Release Notesを変更単位へ分解 - [ ] 日本語タイトル・要約 - [ ] category判定 - [ ] impact判定 - [ ] breaking判定 - [ ] migration判定 - [ ] Gemini API失敗時のリトライ - [ ] APIキーをGitHub Secretsで管理 - [ ] Phase 2の自動収集との統合 - [ ] テスト ## 出力 `changes`, `category`, `title`, `summary`, `impact`, `breaking`, `migration` を構造化JSONとして保存する。 ## 完了条件 - Re…”
haruki33 · GitHub · Aug 17, 2026 - 10
“## 概要 Vertex AI(Gemini Enterprise Agent Platform)の `gemini-2.5-flash` が 2026-10-20 に廃止(ELA 入り)予定のため、Gemini 3 系への移行先モデルを実測で選定し、ADR にまとめる。 Google からの通知では影響プロジェクトとして `documentaisample-488504` が名指しされており、本リポジトリの Gemini 抽出経路が対象。 - 2026-10-20: ELA 入り(呼び出しは継続可能・価格据え置き) - 2027-01-28: ELA 標準価格終了、大幅値上げ+リージョン提供状況の変更可能性 呼び出しが即停止するわけではないため障害リスクは低いが、ADR-0010 が記録したコスト前提(`thinkingBudget:0` での平均 $0.00175/枚)が失効するため、実測をやり直す必要がある。 ## 変更内容 本 Issue のスコープは **移行先モデルの実測比較と方針決定(ADR)まで**。既定モデルの切替コードは別 Issue で対応する。 1. 移行先候…”
git-berian · GitHub · Aug 7, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.