Recommendation for Writing

Writing

Our top recommendation for Writing, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2][3] Deploy for instruction-following writing tasks like paraphrasing, simplifying, and summarization where it scores 75.77% on LiveBench. Watch out: Verify iterative refinement workflows yourself, as one user reported disappointing results when asking for improvements on creative outputs. Google: Gemini 3.6 Flash is the next-ranked alternative. Select for text-heavy workflows where independent testing showed it outperforming DeepSeek v4 Flash Pro.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
9
Revision
v79

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

22

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

2 of 5

Anthropic

Provisional source breadth. 3 citation families and 1 practitioner families support the top result; 1 cautionary thread is retained. The largest citation family contributes 42%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
LMArena Creative Writing
30%
#422/22
LiveBench Instruction Following
25%
#522/22
LMArena Text
20%
#122/22
LiveBench Language
15%
#122/22
OpenRouter usage
10%
79/10022/22

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic40%
  • Anthropic2 models
  • Google1 model
  • Qwen1 model
  • Z.ai1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
84
100%2 threads · 1 families · 1 cautions#1 LiveBench Language · #1 LMArena Text
02Gemini 3.6 FlashGoogle
79
100%1 threads · 1 families · 0 cautions#8 LiveBench Instruction Following · #10 LMArena Creative Writing
03Claude Opus 4.6Anthropic
79
100%1 threads · 1 families · 0 cautions#2 LMArena Creative Writing · #2 LMArena Text
04Qwen3.7 MaxQwen
75
100%no linked practitioner threads#11 LiveBench Instruction Following · #19 LMArena Creative Writing
05GLM 5.2Z.ai
73
100%5 threads · 1 families · 4 cautions#12 LMArena Creative Writing · #24 LMArena Text

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 ranks #4 on LMArena creative writing and scores well on instruction following tasks including story generation, though one user found its iterative refinement disappointing in past testing.

    Best when: Deploy for instruction-following writing tasks like paraphrasing, simplifying, and summarization where it scores 75.77% on LiveBench.

    Tips

    • Deploy for instruction-following writing tasks like paraphrasing, simplifying, and summarization where it scores 75.77% on LiveBench.
      Source 1
      “Scores 75.77% on LiveBench Instruction Following (#5 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗

    Watch out for

    • Verify iterative refinement workflows yourself, as one user reported disappointing results when asking for improvements on creative outputs.
      Source 2
      “I haven't tried this in a few months, but last time I tried a loop that rendered the pelican and asked for improvements the results were actually quite disappointing. Be interesting to try that again against GPT-5.6 at Claude Fable 5 though.”
  2. Gemini 3.6 Flash places #10 on LMArena creative writing with competitive instruction-following scores, and one tester found it superior to DeepSeek v4 Flash Pro on text tasks.

    Best when: Select for text-heavy workflows where independent testing showed it outperforming DeepSeek v4 Flash Pro.

    Tips

    • Select for text-heavy workflows where independent testing showed it outperforming DeepSeek v4 Flash Pro.
      Source 3
      “From my own testing, Gemini 3.5 3.6 Flash is better than DS v4 Flash Pro on text ability.”
      jklmnopqrstuvwOpen original ↗
    • Use for instruction-following writing at 75.37% LiveBench accuracy, covering paraphrasing, story generation, and summarization.
      Source 4
      “Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗
  3. Claude Opus 4.6 sits at #2 on LMArena's creative-writing leaderboard with the highest Elo score among all candidates, indicating strong human preference for its prose quality.

    Best when: Use for premium creative writing where human judges consistently preferred its output over nearly all competitors in blind head-to-head voting.

    Tips

    • Use for premium creative writing where human judges consistently preferred its output over nearly all competitors in blind head-to-head voting.
      Source 5
      “Ranks #2 of 146 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.”
      LMArena creative-writing categoryOpen original ↗
  4. Qwen3.7 Max ranks #20 on LMArena creative writing with solid instruction-following scores, available through Alibaba's platform and several third-party hosts.

    Best when: Use for paraphrasing and story generation where it scores 74.04% on LiveBench instruction following.

    Tips

    • Use for paraphrasing and story generation where it scores 74.04% on LiveBench instruction following.
      Source 6
      “Scores 74.04% on LiveBench Instruction Following (#12 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗
  5. GLM 5.2 is an open-weight model with mixed real-world feedback: one user found it cost-effective for proofreading.

    Best when: Use for proofreading workflows where it outperformed Sonnet 5 on quality and cost in a multi-pass agent benchmark.

    Tips

    • Use for proofreading workflows where it outperformed Sonnet 5 on quality and cost in a multi-pass agent benchmark.
      Source 7
      “I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench”

    Watch out for

    • Expect to manually catch subtle errors in rewritten text, as one user found it introduced mistakes that Sonnet 4.6 corrected.
      Source 8
      “I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…”
    • Verify implementation details yourself, as it generated outdated library versions in coding tasks despite detailed prompts.
      Source 9
      “I don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , di…”
      AgentMasterRaceOpen original ↗

Frequently asked

What is the top-ranked model for Writing?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Deploy for instruction-following writing tasks like paraphrasing, simplifying, and summarization where it scores 75.77% on LiveBench.[1]
What should I watch out for with Anthropic: Claude Fable 5?
Verify iterative refinement workflows yourself, as one user reported disappointing results when asking for improvements on creative outputs.[2]
What is an alternative to Anthropic: Claude Fable 5?
Google: Gemini 3.6 Flash is the next-ranked option. Select for text-heavy workflows where independent testing showed it outperforming DeepSeek v4 Flash Pro.[3]

Sources

  1. 1

    “Scores 75.77% on LiveBench Instruction Following (#5 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  2. 2

    “I haven't tried this in a few months, but last time I tried a loop that rendered the pelican and asked for improvements the results were actually quite disappointing. Be interesting to try that again against GPT-5.6 at Claude Fable 5 though.”

    simonw · Hacker News · Jul 9, 2026
  3. 3

    “From my own testing, Gemini 3.5 3.6 Flash is better than DS v4 Flash Pro on text ability.”

    jklmnopqrstuvw · Hacker News · Aug 13, 2026
  4. 4

    “Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  5. 5

    “Ranks #2 of 146 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.”

    LMArena creative-writing category · Benchmark · Sep 13, 2026
  6. 6

    “Scores 74.04% on LiveBench Instruction Following (#12 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  7. 7

    “I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench”

    artursapek · Hacker News · Jun 30, 2026
  8. 8

    “I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…”

    sixtyj · Hacker News · Jun 30, 2026
  9. 9

    “I don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , di…”

    AgentMasterRace · Hacker News · Jul 7, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.