Recommendation for Long documents

Document Summarization

Our top recommendation for Document Summarization, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use for summarizing lengthy documents where blind human preference data shows reliable handling of long queries, with Elo 1509 placing it in the top tier of 146 models tested. Anthropic: Claude Opus 4.8 is the next-ranked alternative. Use for document summarization when you need Anthropic's Opus-tier reasoning with acceptable long-query performance, ranking #16 of 146 with Elo 1482.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
3
Revision
v76

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

6

task-weighted

Winner coverage

44%

intended feed weight

Largest provider share

3 of 4

Anthropic

Provisional source breadth. 1 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 100%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
FACTS Groundingunavailable
25%
feed unavailable0/20
LiveBench Instruction Following
20%
#517/20
LMArena Document
20%
#412/20
LongBench v2unavailable
15%
feed unavailable0/20
LMArena Long Query
10%
#420/20
OpenRouter usage
10%
79/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic75%
  • Anthropic3 models
  • Google1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
53
44%no linked practitioner threads#4 LMArena Document · #4 LMArena Long Query
02Claude Opus 4.8Anthropic
51
44%no linked practitioner threads#8 LMArena Document · #14 LMArena Long Query
03Gemini 3.6 FlashGoogle
49
44%no linked practitioner threads#8 LiveBench Instruction Following · #17 LMArena Document
04Claude Opus 4.6Anthropic
49
44%no linked practitioner threads#2 LMArena Long Query · #3 LMArena Document

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 ranks fourth on LMArena's long-query category, demonstrating strong performance on extended prompts relevant to document summarization tasks.

    Best when: Use for summarizing lengthy documents where blind human preference data shows reliable handling of long queries, with Elo 1509 placing it in the top tier of 146 models tested.

    Tips

    • Use for summarizing lengthy documents where blind human preference data shows reliable handling of long queries, with Elo 1509 placing it in the top tier of 146 models tested.
      Source 1
      “Ranks #4 of 146 on LMArena's long-query category (Elo 1509), based on blind human preference for longer prompts.”
      LMArena long-query categoryOpen original ↗
  2. Claude Opus 4.8 sits mid-pack on LMArena's long-query leaderboard, with human preference scores suggesting adequate but not exceptional handling of extended prompts.

    Best when: Use for document summarization when you need Anthropic's Opus-tier reasoning with acceptable long-query performance, ranking #16 of 146 with Elo 1482.

    Tips

    • Use for document summarization when you need Anthropic's Opus-tier reasoning with acceptable long-query performance, ranking #16 of 146 with Elo 1482.
      Source 2
      “Ranks #16 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”
      LMArena long-query categoryOpen original ↗
  3. Gemini 3.6 Flash's evaluation comes from a narrower document-specific arena rather than the general long-query benchmark used for other candidates.

    Best when: Use for document tasks specifically, as its #19 ranking on LMArena's document arena with score 1456 reflects direct testing on document-oriented workflows.

    Tips

    • Use for document tasks specifically, as its #19 ranking on LMArena's document arena with score 1456 reflects direct testing on document-oriented workflows.
      Source 3
      “Ranks #19 of 34 on LMArena's document arena (score 1456), based on blind preference for document tasks.”
      LMArena document arenaOpen original ↗

    Watch out for

    • Compare cautiously to other candidates, as the document arena uses a different scoring scale and smaller pool (34 models) than the long-query category, making direct Elo comparisons invalid.
      Source 3
      “Ranks #19 of 34 on LMArena's document arena (score 1456), based on blind preference for document tasks.”
      LMArena document arenaOpen original ↗
  4. Anthropic: Claude Opus 4.6 ranks #2 of 146 on LMArena's long-query category (Elo 1520), based on blind human preference for longer prompts.

    Best when: Consider only after reviewing the cited caution.

Frequently asked

What is the top-ranked model for Document Summarization?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for summarizing lengthy documents where blind human preference data shows reliable handling of long queries, with Elo 1509 placing it in the top tier of 146 models tested.[1]
What is an alternative to Anthropic: Claude Fable 5?
Anthropic: Claude Opus 4.8 is the next-ranked option. Use for document summarization when you need Anthropic's Opus-tier reasoning with acceptable long-query performance, ranking #16 of 146 with Elo 1482.[2]

Sources

  1. 1

    “Ranks #4 of 146 on LMArena's long-query category (Elo 1509), based on blind human preference for longer prompts.”

    LMArena long-query category · Benchmark · Sep 13, 2026
  2. 2

    “Ranks #16 of 146 on LMArena's long-query category (Elo 1482), based on blind human preference for longer prompts.”

    LMArena long-query category · Benchmark · Sep 13, 2026
  3. 3

    “Ranks #19 of 34 on LMArena's document arena (score 1456), based on blind preference for document tasks.”

    LMArena document arena · Benchmark · Sep 13, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.