Recommendation for Long documents
Document Summarization
Our top recommendation for Document Summarization, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use for summarizing lengthy documents where blind human preference data shows reliable handling of long queries, with Elo 1509 placing it in the top tier of 146 models tested. Anthropic: Claude Opus 4.8 is the next-ranked alternative. Use for document summarization when you need Anthropic's Opus-tier reasoning with acceptable long-query performance, ranking #16 of 146 with Elo 1482.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 3
- Revision
- v76
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
6
task-weighted
Winner coverage
44%
intended feed weight
Largest provider share
3 of 4
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| FACTS Groundingunavailable | 25% | feed unavailable | 0/20 |
| LiveBench Instruction Following | 20% | #5 | 17/20 |
| LMArena Document | 20% | #4 | 12/20 |
| LongBench v2unavailable | 15% | feed unavailable | 0/20 |
| LMArena Long Query | 10% | #4 | 20/20 |
| OpenRouter usage | 10% | 79/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic3 models
- Google1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 53 | 44% | no linked practitioner threads | #4 LMArena Document · #4 LMArena Long Query |
| 02 | Claude Opus 4.8Anthropic | 51 | 44% | no linked practitioner threads | #8 LMArena Document · #14 LMArena Long Query |
| 03 | Gemini 3.6 FlashGoogle | 49 | 44% | no linked practitioner threads | #8 LiveBench Instruction Following · #17 LMArena Document |
| 04 | Claude Opus 4.6Anthropic | 49 | 44% | no linked practitioner threads | #2 LMArena Long Query · #3 LMArena Document |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Fable 5 ranks fourth on LMArena's long-query category, demonstrating strong performance on extended prompts relevant to document summarization tasks.
Best when: Use for summarizing lengthy documents where blind human preference data shows reliable handling of long queries, with Elo 1509 placing it in the top tier of 146 models tested.
Tips
- Use for summarizing lengthy documents where blind human preference data shows reliable handling of long queries, with Elo 1509 placing it in the top tier of 146 models tested.
Claude Opus 4.8 sits mid-pack on LMArena's long-query leaderboard, with human preference scores suggesting adequate but not exceptional handling of extended prompts.
Best when: Use for document summarization when you need Anthropic's Opus-tier reasoning with acceptable long-query performance, ranking #16 of 146 with Elo 1482.
Tips
- Use for document summarization when you need Anthropic's Opus-tier reasoning with acceptable long-query performance, ranking #16 of 146 with Elo 1482.
Gemini 3.6 Flash's evaluation comes from a narrower document-specific arena rather than the general long-query benchmark used for other candidates.
Best when: Use for document tasks specifically, as its #19 ranking on LMArena's document arena with score 1456 reflects direct testing on document-oriented workflows.
Tips
- Use for document tasks specifically, as its #19 ranking on LMArena's document arena with score 1456 reflects direct testing on document-oriented workflows.
Watch out for
- Compare cautiously to other candidates, as the document arena uses a different scoring scale and smaller pool (34 models) than the long-query category, making direct Elo comparisons invalid.
Anthropic: Claude Opus 4.6 ranks #2 of 146 on LMArena's long-query category (Elo 1520), based on blind human preference for longer prompts.
Best when: Consider only after reviewing the cited caution.
Frequently asked
- What is the top-ranked model for Document Summarization?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for summarizing lengthy documents where blind human preference data shows reliable handling of long queries, with Elo 1509 placing it in the top tier of 146 models tested.[1]
- What is an alternative to Anthropic: Claude Fable 5?
- Anthropic: Claude Opus 4.8 is the next-ranked option. Use for document summarization when you need Anthropic's Opus-tier reasoning with acceptable long-query performance, ranking #16 of 146 with Elo 1482.[2]
Sources
- 1
“Ranks #4 of 146 on LMArena's long-query category (Elo 1509), based on blind human preference for longer prompts.”
LMArena long-query category · Benchmark · Sep 13, 2026 - 2
“Ranks #16 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”
LMArena long-query category · Benchmark · Sep 13, 2026 - 3
“Ranks #19 of 34 on LMArena's document arena (score 1456), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Sep 13, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.