Recommendation for Chat / Roleplay
Chat & Roleplay
Our top recommendation for Chat & Roleplay, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1][2] Use for high-stakes creative writing and roleplay where human judges consistently preferred its output over nearly all competitors. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Consider if you already have OpenAI infrastructure and need creative writing capability, though human judges preferred 15 other models.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 6
- Revision
- v79
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
1 of 5
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LMArena Creative Writing | 35% | #2 | 21/21 |
| LMArena Text | 25% | #2 | 21/21 |
| LiveBench Instruction Following | 20% | #38 | 21/21 |
| LiveBench Language | 10% | #15 | 21/21 |
| OpenRouter usage | 10% | 83/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic1 model
- Google1 model
- Meta1 model
- OpenAI1 model
- Z.ai1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Opus 4.6Anthropic | 82 | 100% | no linked practitioner threads | #2 LMArena Creative Writing · #2 LMArena Text |
| 02 | GPT-5.6 SolOpenAI | 80 | 100% | no linked practitioner threads | #6 LiveBench Language · #13 LMArena Text |
| 03 | Gemini 3.5 Flash LiteGoogle | 79 | 100% | 1 threads · 1 families · 0 cautions | #27 LiveBench Instruction Following · #34 LMArena Text |
| 04 | Muse Spark 1.2Meta | 77 | 100% | no linked practitioner threads | #4 LMArena Text · #10 LiveBench Instruction Following |
| 05 | GLM 5.2Z.ai | 76 | 100% | 2 threads · 1 families · 2 cautions | #12 LMArena Creative Writing · #24 LMArena Text |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Opus 4.6 holds the second-highest Elo rating in LMArena's creative-writing category among all evaluated models, indicating strong human preference for its conversational and creative outputs.
Best when: Use for high-stakes creative writing and roleplay where human judges consistently preferred its output over nearly all competitors.
Tips
- Use for high-stakes creative writing and roleplay where human judges consistently preferred its output over nearly all competitors.
GPT-5.6 Sol ranks 16th in LMArena's creative-writing category, underperforming relative to its typical benchmark dominance and trailing Claude variants significantly.
Best when: Consider if you already have OpenAI infrastructure and need creative writing capability, though human judges preferred 15 other models.
Tips
- Consider if you already have OpenAI infrastructure and need creative writing capability, though human judges preferred 15 other models.
Watch out for
- Expect lower human preference for creative outputs compared to top-ranked alternatives; this version lags Claude Opus 4.6 by 53 Elo points in blind voting.
Gemini 3.5 Flash Lite has no direct creative-writing benchmark evidence, with only a reported Openrouter connection stability issue in multi-profile workflows.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Watch for Openrouter session handling bugs when switching between connection profiles; one user reported disconnects when moving from summarization to roleplay workflows.
Muse Spark 1.2 ranks 22nd in LMArena's creative-writing category, placing it in the lower half of evaluated models for human preference in conversational and creative tasks.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Expect less compelling creative outputs than most competitors; human judges ranked 21 models higher in blind preference testing.
GLM 5.2 is an open-weight model ranking 12th in LMArena's creative-writing category, though local deployment via Ollama suffers from aggressive quantization and short default context windows.
Best when: Use as an open-weight alternative for creative writing when you can avoid Ollama's 4-bit quants and short context defaults, or deploy through hosted APIs like Fireworks or Together AI.
Tips
- Use as an open-weight alternative for creative writing when you can avoid Ollama's 4-bit quants and short context defaults, or deploy through hosted APIs like Fireworks or Together AI.
Watch out for
- Avoid Ollama deployments for complex multi-turn roleplay; the default 4-bit quantization and short context window breaks coherence on anything beyond simple chat.
Frequently asked
- What is the top-ranked model for Chat & Roleplay?
- Anthropic: Claude Opus 4.6 ranks first in the current evidence-weighted comparison. Use for high-stakes creative writing and roleplay where human judges consistently preferred its output over nearly all competitors.[1]
- What is an alternative to Anthropic: Claude Opus 4.6?
- OpenAI: GPT-5.6 Sol is the next-ranked option. Consider if you already have OpenAI infrastructure and need creative writing capability, though human judges preferred 15 other models.[2]
Sources
- 1
“Ranks #2 of 146 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.”
LMArena creative-writing category · Benchmark · Sep 13, 2026 - 2
“Ranks #16 of 146 on LMArena's creative-writing category (Elo 1452), based on blind human preference votes.”
LMArena creative-writing category · Benchmark · Sep 13, 2026 - 3
“I created a connection profile that used Openrouter as a provider to use Gemini 3.5 flash lite for summarization, and then I would switch to my main roleplay connection profile which used Claude 4.6. It seems that switch disconnects the profile from Openrouter. So, when summarization activated, it would fail.”
Cheesedozer · GitHub · Aug 2, 2026 - 4
“Ranks #22 of 146 on LMArena's creative-writing category (Elo 1448), based on blind human preference votes.”
LMArena creative-writing category · Benchmark · Sep 13, 2026 - 5
“Ranks #12 of 146 on LMArena's creative-writing category (Elo 1461), based on blind human preference votes.”
LMArena creative-writing category · Benchmark · Sep 13, 2026 - 6
“Ollama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.”
kgeist · Hacker News · Jul 13, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.