Recommendation for Summarization

Summarization

Our top recommendation for Summarization, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2][3] Use for long-document summarization where human preference rankings matter, as it scores 1509 Elo on LMArena's long-query benchmark. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Consider for agentic chat pipelines where you need summarization alongside other instruction-following tasks, though verify latency against 5.4 variants first.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
11
Revision
v78

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

6

task-weighted

Winner coverage

44%

intended feed weight

Largest provider share

3 of 8

Anthropic

Provisional source breadth. 3 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 46%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
FACTS Groundingunavailable
25%
feed unavailable0/20
LiveBench Instruction Following
20%
#517/20
LMArena Document
20%
#412/20
LongBench v2unavailable
15%
feed unavailable0/20
LMArena Long Query
10%
#420/20
OpenRouter usage
10%
79/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic38%
  • Anthropic3 models
  • OpenAI2 models
  • Google1 model
  • Qwen1 model
  • xiaomi1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
53
44%no linked practitioner threads#4 LMArena Document · #4 LMArena Long Query
02GPT-5.6 SolOpenAI
52
44%1 threads · 1 families · 0 cautions#6 LMArena Document · #17 LiveBench Instruction Following
03GPT-6 AstraOpenAI
52
44%1 threads · 1 families · 0 cautions#7 LiveBench Instruction Following · #11 LMArena Document
04Claude Opus 4.8Anthropic
51
44%no linked practitioner threads#8 LMArena Document · #14 LMArena Long Query
05Gemini 3.6 FlashGoogle
49
44%no linked practitioner threads#8 LiveBench Instruction Following · #17 LMArena Document
06Claude Opus 4.6Anthropic
49
44%no linked practitioner threads#2 LMArena Long Query · #3 LMArena Document
07Qwen3.8 27BQwen
49
38%no linked practitioner threads#14 LiveBench Instruction Following · #37 LMArena Long Query
08MiMo-V2.5-Proxiaomi
49
28%no linked practitioner threads#14 LMArena Long Query

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 places #4 on LMArena's long-query category and scores 75.77% on LiveBench Instruction Following, which includes summarization tasks.

    Best when: Use for long-document summarization where human preference rankings matter, as it scores 1509 Elo on LMArena's long-query benchmark.

    Tips

    • Use for long-document summarization where human preference rankings matter, as it scores 1509 Elo on LMArena's long-query benchmark.
      Source 1
      “Ranks #4 of 146 on LMArena's long-query category (Elo 1509), based on blind human preference for longer prompts.”
      LMArena long-query categoryOpen original ↗
    • Deploy for instruction-following workflows that mix summarization with paraphrasing and simplification, given its 75.77% LiveBench score.
      Source 4
      “Scores 75.77% on LiveBench Instruction Following (#5 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗
  2. GPT-5.6 Sol ranks #20 on LMArena's long-query category and scores 71.85% on LiveBench Instruction Following, with community reports noting 5.4-mini remains preferred for quick summaries due to speed.

    Best when: Consider for agentic chat pipelines where you need summarization alongside other instruction-following tasks, though verify latency against 5.4 variants first.

    Tips

    • Consider for agentic chat pipelines where you need summarization alongside other instruction-following tasks, though verify latency against 5.4 variants first.
      Source 2
      “I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?”
      Source 3
      “Scores 71.85% on LiveBench Instruction Following (#18 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗

    Watch out for

    • Benchmark before replacing GPT-5.4-mini for quick summaries and TLDRs, as practitioners report 5.4 delivers comparable results with faster throughput.
      Source 2
      “I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?”
  3. GPT-6 Astra scores 75.58% on LiveBench Instruction Following, placing #7 of 58, which includes summarization among its evaluated tasks.

    Best when: Use for mixed instruction-following workloads that combine summarization with paraphrasing and story generation, where its 75.58% score indicates strong capability.

    Tips

    • Use for mixed instruction-following workloads that combine summarization with paraphrasing and story generation, where its 75.58% score indicates strong capability.
      Source 5
      “Scores 75.58% on LiveBench Instruction Following (#7 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗

    Watch out for

    • No LMArena long-query or document-specific ranking is available to confirm performance on extended context summarization.
      Source 5
      “Scores 75.58% on LiveBench Instruction Following (#7 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗
  4. Claude Opus 4.8 ranks #16 on LMArena's long-query category and scores 72.03% on LiveBench Instruction Following, including summarization tasks.

    Best when: Deploy for long-context summarization where its 1482 Elo on LMArena long-query indicates solid human preference performance.

    Tips

    • Deploy for long-context summarization where its 1482 Elo on LMArena long-query indicates solid human preference performance.
      Source 6
      “Ranks #16 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”
      LMArena long-query categoryOpen original ↗
  5. Gemini 3.6 Flash ranks #19 on LMArena's document arena and scores 75.37% on LiveBench Instruction Following, with document-specific benchmarking.

    Best when: Use for document-focused summarization tasks where its 1456 score on LMArena's document-specific arena provides direct relevance signal.

    Tips

    • Use for document-focused summarization tasks where its 1456 score on LMArena's document-specific arena provides direct relevance signal.
      Source 7
      “Ranks #19 of 34 on LMArena's document arena (score 1456), based on blind preference for document tasks.”
      LMArena document arenaOpen original ↗
    • Deploy for cost-sensitive summarization pipelines that mix summarization with simplification and paraphrasing, given its 75.37% LiveBench Instruction Following score.
      Source 8
      “Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗

    Watch out for

    • Its #19 ranking on the document arena trails top performers, and no long-query specific ranking is available for extended context evaluation.
      Source 7
      “Ranks #19 of 34 on LMArena's document arena (score 1456), based on blind preference for document tasks.”
      LMArena document arenaOpen original ↗
  6. Claude Opus 4.6 ranks #2 on LMArena's long-query category with 1520 Elo, the highest long-query preference score among all candidates.

    Best when: Prioritize for long-document and extended-thread summarization where human preference rankings are critical, as its #2 position indicates top-tier output quality.

    Tips

    • Prioritize for long-document and extended-thread summarization where human preference rankings are critical, as its #2 position indicates top-tier output quality.
      Source 9
      “Ranks #2 of 146 on LMArena's long-query category (Elo 1520), based on blind human preference for longer prompts.”
      LMArena long-query categoryOpen original ↗

    Watch out for

    • No LiveBench Instruction Following score is available to verify performance on instruction-heavy summarization workflows that mix tasks.
      Source 9
      “Ranks #2 of 146 on LMArena's long-query category (Elo 1520), based on blind human preference for longer prompts.”
      LMArena long-query categoryOpen original ↗
  7. Qwen3.8 27B scores 72.66% on LiveBench Instruction Following, placing #15 of 58, which includes summarization among evaluated tasks.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • No LMArena long-query or document-specific ranking is available to assess performance on extended context summarization.
      Source 10
      “Scores 72.66% on LiveBench Instruction Following (#15 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗
  8. MiMo-V2.5-Pro ranks #15 on LMArena's long-query category with 1482 Elo, based on blind human preference for longer prompts.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • No LiveBench Instruction Following score is available to verify capability on instruction-heavy summarization that mixes paraphrasing or simplification.
      Source 11
      “Ranks #15 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”
      LMArena long-query categoryOpen original ↗

Frequently asked

What is the top-ranked model for Summarization?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for long-document summarization where human preference rankings matter, as it scores 1509 Elo on LMArena's long-query benchmark.[1]
What is an alternative to Anthropic: Claude Fable 5?
OpenAI: GPT-5.6 Sol is the next-ranked option. Consider for agentic chat pipelines where you need summarization alongside other instruction-following tasks, though verify latency against 5.4 variants first.[2][3]

Sources

  1. 1

    “Ranks #4 of 146 on LMArena's long-query category (Elo 1509), based on blind human preference for longer prompts.”

    LMArena long-query category · Benchmark · Sep 13, 2026
  2. 2

    “I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?”

    thomas_witt · Hacker News · Jul 10, 2026
  3. 3

    “Scores 71.85% on LiveBench Instruction Following (#18 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  4. 4

    “Scores 75.77% on LiveBench Instruction Following (#5 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  5. 5

    “Scores 75.58% on LiveBench Instruction Following (#7 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  6. 6

    “Ranks #16 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”

    LMArena long-query category · Benchmark · Sep 13, 2026
  7. 7

    “Ranks #19 of 34 on LMArena's document arena (score 1456), based on blind preference for document tasks.”

    LMArena document arena · Benchmark · Sep 13, 2026
  8. 8

    “Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  9. 9

    “Ranks #2 of 146 on LMArena's long-query category (Elo 1520), based on blind human preference for longer prompts.”

    LMArena long-query category · Benchmark · Sep 13, 2026
  10. 10

    “Scores 72.66% on LiveBench Instruction Following (#15 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  11. 11

    “Ranks #15 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”

    LMArena long-query category · Benchmark · Sep 13, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.