Recommendation for Summarization
Summarization
Our top recommendation for Summarization, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2][3] Use for long-document summarization where human preference rankings matter, as it scores 1509 Elo on LMArena's long-query benchmark. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Consider for agentic chat pipelines where you need summarization alongside other instruction-following tasks, though verify latency against 5.4 variants first.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 11
- Revision
- v78
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
6
task-weighted
Winner coverage
44%
intended feed weight
Largest provider share
3 of 8
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| FACTS Groundingunavailable | 25% | feed unavailable | 0/20 |
| LiveBench Instruction Following | 20% | #5 | 17/20 |
| LMArena Document | 20% | #4 | 12/20 |
| LongBench v2unavailable | 15% | feed unavailable | 0/20 |
| LMArena Long Query | 10% | #4 | 20/20 |
| OpenRouter usage | 10% | 79/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic3 models
- OpenAI2 models
- Google1 model
- Qwen1 model
- xiaomi1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 53 | 44% | no linked practitioner threads | #4 LMArena Document · #4 LMArena Long Query |
| 02 | GPT-5.6 SolOpenAI | 52 | 44% | 1 threads · 1 families · 0 cautions | #6 LMArena Document · #17 LiveBench Instruction Following |
| 03 | GPT-6 AstraOpenAI | 52 | 44% | 1 threads · 1 families · 0 cautions | #7 LiveBench Instruction Following · #11 LMArena Document |
| 04 | Claude Opus 4.8Anthropic | 51 | 44% | no linked practitioner threads | #8 LMArena Document · #14 LMArena Long Query |
| 05 | Gemini 3.6 FlashGoogle | 49 | 44% | no linked practitioner threads | #8 LiveBench Instruction Following · #17 LMArena Document |
| 06 | Claude Opus 4.6Anthropic | 49 | 44% | no linked practitioner threads | #2 LMArena Long Query · #3 LMArena Document |
| 07 | Qwen3.8 27BQwen | 49 | 38% | no linked practitioner threads | #14 LiveBench Instruction Following · #37 LMArena Long Query |
| 08 | MiMo-V2.5-Proxiaomi | 49 | 28% | no linked practitioner threads | #14 LMArena Long Query |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Fable 5 places #4 on LMArena's long-query category and scores 75.77% on LiveBench Instruction Following, which includes summarization tasks.
Best when: Use for long-document summarization where human preference rankings matter, as it scores 1509 Elo on LMArena's long-query benchmark.
Tips
- Use for long-document summarization where human preference rankings matter, as it scores 1509 Elo on LMArena's long-query benchmark.
- Deploy for instruction-following workflows that mix summarization with paraphrasing and simplification, given its 75.77% LiveBench score.
GPT-5.6 Sol ranks #20 on LMArena's long-query category and scores 71.85% on LiveBench Instruction Following, with community reports noting 5.4-mini remains preferred for quick summaries due to speed.
Best when: Consider for agentic chat pipelines where you need summarization alongside other instruction-following tasks, though verify latency against 5.4 variants first.
Tips
- Consider for agentic chat pipelines where you need summarization alongside other instruction-following tasks, though verify latency against 5.4 variants first.
Watch out for
- Benchmark before replacing GPT-5.4-mini for quick summaries and TLDRs, as practitioners report 5.4 delivers comparable results with faster throughput.
GPT-6 Astra scores 75.58% on LiveBench Instruction Following, placing #7 of 58, which includes summarization among its evaluated tasks.
Best when: Use for mixed instruction-following workloads that combine summarization with paraphrasing and story generation, where its 75.58% score indicates strong capability.
Tips
- Use for mixed instruction-following workloads that combine summarization with paraphrasing and story generation, where its 75.58% score indicates strong capability.
Watch out for
- No LMArena long-query or document-specific ranking is available to confirm performance on extended context summarization.
Claude Opus 4.8 ranks #16 on LMArena's long-query category and scores 72.03% on LiveBench Instruction Following, including summarization tasks.
Best when: Deploy for long-context summarization where its 1482 Elo on LMArena long-query indicates solid human preference performance.
Tips
- Deploy for long-context summarization where its 1482 Elo on LMArena long-query indicates solid human preference performance.
Gemini 3.6 Flash ranks #19 on LMArena's document arena and scores 75.37% on LiveBench Instruction Following, with document-specific benchmarking.
Best when: Use for document-focused summarization tasks where its 1456 score on LMArena's document-specific arena provides direct relevance signal.
Tips
- Use for document-focused summarization tasks where its 1456 score on LMArena's document-specific arena provides direct relevance signal.
- Deploy for cost-sensitive summarization pipelines that mix summarization with simplification and paraphrasing, given its 75.37% LiveBench Instruction Following score.
Watch out for
- Its #19 ranking on the document arena trails top performers, and no long-query specific ranking is available for extended context evaluation.
Claude Opus 4.6 ranks #2 on LMArena's long-query category with 1520 Elo, the highest long-query preference score among all candidates.
Best when: Prioritize for long-document and extended-thread summarization where human preference rankings are critical, as its #2 position indicates top-tier output quality.
Tips
- Prioritize for long-document and extended-thread summarization where human preference rankings are critical, as its #2 position indicates top-tier output quality.
Watch out for
- No LiveBench Instruction Following score is available to verify performance on instruction-heavy summarization workflows that mix tasks.
Qwen3.8 27B scores 72.66% on LiveBench Instruction Following, placing #15 of 58, which includes summarization among evaluated tasks.
Best when: Consider only after reviewing the cited caution.
Watch out for
- No LMArena long-query or document-specific ranking is available to assess performance on extended context summarization.
MiMo-V2.5-Pro ranks #15 on LMArena's long-query category with 1482 Elo, based on blind human preference for longer prompts.
Best when: Consider only after reviewing the cited caution.
Watch out for
- No LiveBench Instruction Following score is available to verify capability on instruction-heavy summarization that mixes paraphrasing or simplification.
Frequently asked
- What is the top-ranked model for Summarization?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for long-document summarization where human preference rankings matter, as it scores 1509 Elo on LMArena's long-query benchmark.[1]
Sources
- 1
“Ranks #4 of 146 on LMArena's long-query category (Elo 1509), based on blind human preference for longer prompts.”
LMArena long-query category · Benchmark · Sep 13, 2026 - 2
“I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?”
thomas_witt · Hacker News · Jul 10, 2026 - 3
“Scores 71.85% on LiveBench Instruction Following (#18 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 4
“Scores 75.77% on LiveBench Instruction Following (#5 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 5
“Scores 75.58% on LiveBench Instruction Following (#7 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 6
“Ranks #16 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”
LMArena long-query category · Benchmark · Sep 13, 2026 - 7
“Ranks #19 of 34 on LMArena's document arena (score 1456), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Sep 13, 2026 - 8
“Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 9
“Ranks #2 of 146 on LMArena's long-query category (Elo 1520), based on blind human preference for longer prompts.”
LMArena long-query category · Benchmark · Sep 13, 2026 - 10
“Scores 72.66% on LiveBench Instruction Following (#15 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 11
“Ranks #15 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”
LMArena long-query category · Benchmark · Sep 13, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.