Recommendation for Business & professional

Business Writing

Our top recommendation for Business Writing, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1][2] Use when you need documents that feel natural and well-organized to human readers, as it tops LMArena's instruction-following category with the highest Elo score (1523) in blind head-to-head voting. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Choose over reasoning-heavy variants when you need clear, controlled prose in collaborative documents, as one user abandoned a high-reasoning model for duplicating lines and unauthorized edits, returning to this for reliable writing.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
12
Revision
v77

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

1 of 6

Anthropic

Provisional source breadth. 3 citation families and 1 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 46%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Opus 4.6
Evaluation feedWeightWinner resultField measured
LiveBench Instruction Following
30%
#3820/21
LMArena Instruction Following
25%
#121/21
LMArena Text
20%
#221/21
LMArena Creative Writing
15%
#221/21
OpenRouter usage
10%
83/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic17%
  • Anthropic1 model
  • Google1 model
  • Meta1 model
  • OpenAI1 model
  • xiaomi1 model
  • Z.ai1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Opus 4.6Anthropic
80
100%1 threads · 1 families · 0 cautions#1 LMArena Instruction Following · #2 LMArena Creative Writing
02GPT-5.6 SolOpenAI
80
100%1 threads · 1 families · 0 cautions#12 LMArena Instruction Following · #13 LMArena Text
03Gemini 3.6 FlashGoogle
80
100%no linked practitioner threads#8 LiveBench Instruction Following · #10 LMArena Creative Writing
04Muse Spark 1.2Meta
77
100%no linked practitioner threads#4 LMArena Text · #10 LiveBench Instruction Following
05GLM 5.2Z.ai
74
100%3 threads · 1 families · 3 cautions#12 LMArena Creative Writing · #21 LMArena Instruction Following
06MiMo-V2.5-Proxiaomi
74
87%no linked practitioner threads#10 LMArena Instruction Following · #26 LMArena Text

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Leads human preference rankings for instruction following, suggesting strong alignment with how users actually want business documents structured and phrased.

    Best when: Use when you need documents that feel natural and well-organized to human readers, as it tops LMArena's instruction-following category with the highest Elo score (1523) in blind head-to-head voting.

    Tips

    • Use when you need documents that feel natural and well-organized to human readers, as it tops LMArena's instruction-following category with the highest Elo score (1523) in blind head-to-head voting.
      Source 1
      “Ranks #1 of 146 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.”
      LMArena instruction-following categoryOpen original ↗
  2. Preferred by at least one user over high-reasoning alternatives for design document clarity, though benchmark rankings place it mid-pack for instruction following.

    Best when: Choose over reasoning-heavy variants when you need clear, controlled prose in collaborative documents, as one user abandoned a high-reasoning model for duplicating lines and unauthorized edits, returning to this for reliable writing.

    Tips

    • Choose over reasoning-heavy variants when you need clear, controlled prose in collaborative documents, as one user abandoned a high-reasoning model for duplicating lines and unauthorized edits, returning to this for reliable writing.
      Source 2
      “I tried Astra w high reasoning on a design document project and it was horrible. It started duplicating output lines, made document edits without permission, and basically did a poor job writing clear prose. I went back to 5.6-sol and it's great. I'm an OpenAI fanboy and was severely disappointed. I hope Astra is better for coding.”

    Watch out for

    • Expect middle-tier benchmark performance on instruction following, ranking #13 on LMArena (Elo 1477) and #18 on LiveBench at 71.85%, below several competitors in this list.
      Source 3
      “Ranks #13 of 146 on LMArena's instruction-following category (Elo 1477), based on blind human preference votes.”
      LMArena instruction-following categoryOpen original ↗
      Source 4
      “Scores 71.85% on LiveBench Instruction Following (#18 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗
  3. Competitive LiveBench scores for instruction following including summarization, though human preference rankings place it lower than its benchmark standing would suggest.

    Best when: Use for structured rewriting tasks like paraphrasing and simplification, where it scores 75.37% on LiveBench instruction following (#8 of 58).

    Tips

    • Use for structured rewriting tasks like paraphrasing and simplification, where it scores 75.37% on LiveBench instruction following (#8 of 58).
      Source 5
      “Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗

    Watch out for

    • Note the gap between benchmark and preference performance, ranking only #20 on LMArena (Elo 1467) despite stronger LiveBench showing, which may indicate prose style less aligned with human reviewers.
      Source 6
      “Ranks #20 of 146 on LMArena's instruction-following category (Elo 1467), based on blind human preference votes.”
      LMArena instruction-following categoryOpen original ↗
  4. Solid mid-tier performance on both human preference and structured instruction-following evaluations, with no specific business writing evidence beyond general rankings.

    Best when: Consider for balanced instruction-following tasks where benchmark consistency matters, scoring 74.33% on LiveBench (#10 of 58) and placing #18 on LMArena (Elo 1471).

    Tips

    • Consider for balanced instruction-following tasks where benchmark consistency matters, scoring 74.33% on LiveBench (#10 of 58) and placing #18 on LMArena (Elo 1471).
      Source 7
      “Ranks #18 of 146 on LMArena's instruction-following category (Elo 1471), based on blind human preference votes.”
      LMArena instruction-following categoryOpen original ↗
      Source 8
      “Scores 74.33% on LiveBench Instruction Following (#10 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗
  5. An open-weight model with strong benchmark numbers but mixed real-world feedback, including reports of subtle errors in rewriting tasks that required manual correction.

    Best when: Consider for cost-sensitive proofreading workflows where benchmark results translate, as one independent tester found it superior to Sonnet 5 on quality and cost in an English error-correction benchmark with agent loops.

    Tips

    • Consider for cost-sensitive proofreading workflows where benchmark results translate, as one independent tester found it superior to Sonnet 5 on quality and cost in an English error-correction benchmark with agent loops.
      Source 9
      “I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench”

    Watch out for

    • Plan for additional review cycles on critical documents, as multiple users report subtle mistakes in rewriting and planning tasks that required hand-correction, with one noting it seemed good "on paper" but produced different real usage results.
      Source 10
      “I don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , di…”
      AgentMasterRaceOpen original ↗
      Source 11
      “I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…”
  6. An open-weight option that places in the top dozen for instruction-following human preferences, offering a deployable alternative for organizations with infrastructure constraints.

    Best when: Self-host or route through open-weight providers when you need control over deployment and data residency, while still achieving strong instruction-following performance (#12 on LMArena with Elo 1478).

    Tips

    • Self-host or route through open-weight providers when you need control over deployment and data residency, while still achieving strong instruction-following performance (#12 on LMArena with Elo 1478).
      Source 12
      “Ranks #12 of 146 on LMArena's instruction-following category (Elo 1478), based on blind human preference votes.”
      LMArena instruction-following categoryOpen original ↗

Frequently asked

What is the top-ranked model for Business Writing?
Anthropic: Claude Opus 4.6 ranks first in the current evidence-weighted comparison. Use when you need documents that feel natural and well-organized to human readers, as it tops LMArena's instruction-following category with the highest Elo score (1523) in blind head-to-head voting.[1]
What is an alternative to Anthropic: Claude Opus 4.6?
OpenAI: GPT-5.6 Sol is the next-ranked option. Choose over reasoning-heavy variants when you need clear, controlled prose in collaborative documents, as one user abandoned a high-reasoning model for duplicating lines and unauthorized edits, returning to this for reliable writing.[2]

Sources

  1. 1

    “Ranks #1 of 146 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.”

    LMArena instruction-following category · Benchmark · Sep 13, 2026
  2. 2

    “I tried Astra w high reasoning on a design document project and it was horrible. It started duplicating output lines, made document edits without permission, and basically did a poor job writing clear prose. I went back to 5.6-sol and it's great. I'm an OpenAI fanboy and was severely disappointed. I hope Astra is better for coding.”

    01100011 · Hacker News · Sep 21, 2026
  3. 3

    “Ranks #13 of 146 on LMArena's instruction-following category (Elo 1477), based on blind human preference votes.”

    LMArena instruction-following category · Benchmark · Sep 13, 2026
  4. 4

    “Scores 71.85% on LiveBench Instruction Following (#18 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  5. 5

    “Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  6. 6

    “Ranks #20 of 146 on LMArena's instruction-following category (Elo 1467), based on blind human preference votes.”

    LMArena instruction-following category · Benchmark · Sep 13, 2026
  7. 7

    “Ranks #18 of 146 on LMArena's instruction-following category (Elo 1471), based on blind human preference votes.”

    LMArena instruction-following category · Benchmark · Sep 13, 2026
  8. 8

    “Scores 74.33% on LiveBench Instruction Following (#10 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  9. 9

    “I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench”

    artursapek · Hacker News · Jun 30, 2026
  10. 10

    “I don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , di…”

    AgentMasterRace · Hacker News · Jul 7, 2026
  11. 11

    “I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…”

    sixtyj · Hacker News · Jun 30, 2026
  12. 12

    “Ranks #12 of 146 on LMArena's instruction-following category (Elo 1478), based on blind human preference votes.”

    LMArena instruction-following category · Benchmark · Sep 13, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.