Recommendation for Business & professional
Business Writing
Our top recommendation for Business Writing, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1][2] Use when you need documents that feel natural and well-organized to human readers, as it tops LMArena's instruction-following category with the highest Elo score (1523) in blind head-to-head voting. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Choose over reasoning-heavy variants when you need clear, controlled prose in collaborative documents, as one user abandoned a high-reasoning model for duplicating lines and unauthorized edits, returning to this for reliable writing.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 12
- Revision
- v77
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
1 of 6
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LiveBench Instruction Following | 30% | #38 | 20/21 |
| LMArena Instruction Following | 25% | #1 | 21/21 |
| LMArena Text | 20% | #2 | 21/21 |
| LMArena Creative Writing | 15% | #2 | 21/21 |
| OpenRouter usage | 10% | 83/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic1 model
- Google1 model
- Meta1 model
- OpenAI1 model
- xiaomi1 model
- Z.ai1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Opus 4.6Anthropic | 80 | 100% | 1 threads · 1 families · 0 cautions | #1 LMArena Instruction Following · #2 LMArena Creative Writing |
| 02 | GPT-5.6 SolOpenAI | 80 | 100% | 1 threads · 1 families · 0 cautions | #12 LMArena Instruction Following · #13 LMArena Text |
| 03 | Gemini 3.6 FlashGoogle | 80 | 100% | no linked practitioner threads | #8 LiveBench Instruction Following · #10 LMArena Creative Writing |
| 04 | Muse Spark 1.2Meta | 77 | 100% | no linked practitioner threads | #4 LMArena Text · #10 LiveBench Instruction Following |
| 05 | GLM 5.2Z.ai | 74 | 100% | 3 threads · 1 families · 3 cautions | #12 LMArena Creative Writing · #21 LMArena Instruction Following |
| 06 | MiMo-V2.5-Proxiaomi | 74 | 87% | no linked practitioner threads | #10 LMArena Instruction Following · #26 LMArena Text |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Leads human preference rankings for instruction following, suggesting strong alignment with how users actually want business documents structured and phrased.
Best when: Use when you need documents that feel natural and well-organized to human readers, as it tops LMArena's instruction-following category with the highest Elo score (1523) in blind head-to-head voting.
Tips
- Use when you need documents that feel natural and well-organized to human readers, as it tops LMArena's instruction-following category with the highest Elo score (1523) in blind head-to-head voting.
Preferred by at least one user over high-reasoning alternatives for design document clarity, though benchmark rankings place it mid-pack for instruction following.
Best when: Choose over reasoning-heavy variants when you need clear, controlled prose in collaborative documents, as one user abandoned a high-reasoning model for duplicating lines and unauthorized edits, returning to this for reliable writing.
Tips
- Choose over reasoning-heavy variants when you need clear, controlled prose in collaborative documents, as one user abandoned a high-reasoning model for duplicating lines and unauthorized edits, returning to this for reliable writing.
Watch out for
- Expect middle-tier benchmark performance on instruction following, ranking #13 on LMArena (Elo 1477) and #18 on LiveBench at 71.85%, below several competitors in this list.
Competitive LiveBench scores for instruction following including summarization, though human preference rankings place it lower than its benchmark standing would suggest.
Best when: Use for structured rewriting tasks like paraphrasing and simplification, where it scores 75.37% on LiveBench instruction following (#8 of 58).
Tips
- Use for structured rewriting tasks like paraphrasing and simplification, where it scores 75.37% on LiveBench instruction following (#8 of 58).
Watch out for
- Note the gap between benchmark and preference performance, ranking only #20 on LMArena (Elo 1467) despite stronger LiveBench showing, which may indicate prose style less aligned with human reviewers.
Solid mid-tier performance on both human preference and structured instruction-following evaluations, with no specific business writing evidence beyond general rankings.
Best when: Consider for balanced instruction-following tasks where benchmark consistency matters, scoring 74.33% on LiveBench (#10 of 58) and placing #18 on LMArena (Elo 1471).
Tips
- Consider for balanced instruction-following tasks where benchmark consistency matters, scoring 74.33% on LiveBench (#10 of 58) and placing #18 on LMArena (Elo 1471).
An open-weight model with strong benchmark numbers but mixed real-world feedback, including reports of subtle errors in rewriting tasks that required manual correction.
Best when: Consider for cost-sensitive proofreading workflows where benchmark results translate, as one independent tester found it superior to Sonnet 5 on quality and cost in an English error-correction benchmark with agent loops.
Tips
- Consider for cost-sensitive proofreading workflows where benchmark results translate, as one independent tester found it superior to Sonnet 5 on quality and cost in an English error-correction benchmark with agent loops.
Watch out for
- Plan for additional review cycles on critical documents, as multiple users report subtle mistakes in rewriting and planning tasks that required hand-correction, with one noting it seemed good "on paper" but produced different real usage results.
An open-weight option that places in the top dozen for instruction-following human preferences, offering a deployable alternative for organizations with infrastructure constraints.
Best when: Self-host or route through open-weight providers when you need control over deployment and data residency, while still achieving strong instruction-following performance (#12 on LMArena with Elo 1478).
Tips
- Self-host or route through open-weight providers when you need control over deployment and data residency, while still achieving strong instruction-following performance (#12 on LMArena with Elo 1478).
Frequently asked
- What is the top-ranked model for Business Writing?
- Anthropic: Claude Opus 4.6 ranks first in the current evidence-weighted comparison. Use when you need documents that feel natural and well-organized to human readers, as it tops LMArena's instruction-following category with the highest Elo score (1523) in blind head-to-head voting.[1]
- What is an alternative to Anthropic: Claude Opus 4.6?
- OpenAI: GPT-5.6 Sol is the next-ranked option. Choose over reasoning-heavy variants when you need clear, controlled prose in collaborative documents, as one user abandoned a high-reasoning model for duplicating lines and unauthorized edits, returning to this for reliable writing.[2]
Sources
- 1
“Ranks #1 of 146 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 13, 2026 - 2
“I tried Astra w high reasoning on a design document project and it was horrible. It started duplicating output lines, made document edits without permission, and basically did a poor job writing clear prose. I went back to 5.6-sol and it's great. I'm an OpenAI fanboy and was severely disappointed. I hope Astra is better for coding.”
01100011 · Hacker News · Sep 21, 2026 - 3
“Ranks #13 of 146 on LMArena's instruction-following category (Elo 1477), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 13, 2026 - 4
“Scores 71.85% on LiveBench Instruction Following (#18 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 5
“Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 6
“Ranks #20 of 146 on LMArena's instruction-following category (Elo 1467), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 13, 2026 - 7
“Ranks #18 of 146 on LMArena's instruction-following category (Elo 1471), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 13, 2026 - 8
“Scores 74.33% on LiveBench Instruction Following (#10 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 9
“I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench”
artursapek · Hacker News · Jun 30, 2026 - 10
“I don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , di…”
AgentMasterRace · Hacker News · Jul 7, 2026 - 11
“I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…”
sixtyj · Hacker News · Jun 30, 2026 - 12
“Ranks #12 of 146 on LMArena's instruction-following category (Elo 1478), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 13, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.