Recommendation for Chat / Roleplay

Chat & Roleplay

Our top recommendation for Chat & Roleplay, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1][2] Use for high-stakes creative writing and roleplay where human judges consistently preferred its output over nearly all competitors. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Consider if you already have OpenAI infrastructure and need creative writing capability, though human judges preferred 15 other models.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
6
Revision
v79

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

1 of 5

Anthropic

Provisional source breadth. 3 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 67%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Opus 4.6
Evaluation feedWeightWinner resultField measured
LMArena Creative Writing
35%
#221/21
LMArena Text
25%
#221/21
LiveBench Instruction Following
20%
#3821/21
LiveBench Language
10%
#1521/21
OpenRouter usage
10%
83/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic20%
  • Anthropic1 model
  • Google1 model
  • Meta1 model
  • OpenAI1 model
  • Z.ai1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Opus 4.6Anthropic
82
100%no linked practitioner threads#2 LMArena Creative Writing · #2 LMArena Text
02GPT-5.6 SolOpenAI
80
100%no linked practitioner threads#6 LiveBench Language · #13 LMArena Text
03Gemini 3.5 Flash LiteGoogle
79
100%1 threads · 1 families · 0 cautions#27 LiveBench Instruction Following · #34 LMArena Text
04Muse Spark 1.2Meta
77
100%no linked practitioner threads#4 LMArena Text · #10 LiveBench Instruction Following
05GLM 5.2Z.ai
76
100%2 threads · 1 families · 2 cautions#12 LMArena Creative Writing · #24 LMArena Text

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Opus 4.6 holds the second-highest Elo rating in LMArena's creative-writing category among all evaluated models, indicating strong human preference for its conversational and creative outputs.

    Best when: Use for high-stakes creative writing and roleplay where human judges consistently preferred its output over nearly all competitors.

    Tips

    • Use for high-stakes creative writing and roleplay where human judges consistently preferred its output over nearly all competitors.
      Source 1
      “Ranks #2 of 146 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.”
      LMArena creative-writing categoryOpen original ↗
  2. GPT-5.6 Sol ranks 16th in LMArena's creative-writing category, underperforming relative to its typical benchmark dominance and trailing Claude variants significantly.

    Best when: Consider if you already have OpenAI infrastructure and need creative writing capability, though human judges preferred 15 other models.

    Tips

    • Consider if you already have OpenAI infrastructure and need creative writing capability, though human judges preferred 15 other models.
      Source 2
      “Ranks #16 of 146 on LMArena's creative-writing category (Elo 1452), based on blind human preference votes.”
      LMArena creative-writing categoryOpen original ↗

    Watch out for

    • Expect lower human preference for creative outputs compared to top-ranked alternatives; this version lags Claude Opus 4.6 by 53 Elo points in blind voting.
      Source 2
      “Ranks #16 of 146 on LMArena's creative-writing category (Elo 1452), based on blind human preference votes.”
      LMArena creative-writing categoryOpen original ↗
  3. Gemini 3.5 Flash Lite has no direct creative-writing benchmark evidence, with only a reported Openrouter connection stability issue in multi-profile workflows.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Watch for Openrouter session handling bugs when switching between connection profiles; one user reported disconnects when moving from summarization to roleplay workflows.
      Source 3
      “I created a connection profile that used Openrouter as a provider to use Gemini 3.5 flash lite for summarization, and then I would switch to my main roleplay connection profile which used Claude 4.6. It seems that switch disconnects the profile from Openrouter. So, when summarization activated, it would fail.”
  4. Muse Spark 1.2 ranks 22nd in LMArena's creative-writing category, placing it in the lower half of evaluated models for human preference in conversational and creative tasks.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Expect less compelling creative outputs than most competitors; human judges ranked 21 models higher in blind preference testing.
      Source 4
      “Ranks #22 of 146 on LMArena's creative-writing category (Elo 1448), based on blind human preference votes.”
      LMArena creative-writing categoryOpen original ↗
  5. GLM 5.2 is an open-weight model ranking 12th in LMArena's creative-writing category, though local deployment via Ollama suffers from aggressive quantization and short default context windows.

    Best when: Use as an open-weight alternative for creative writing when you can avoid Ollama's 4-bit quants and short context defaults, or deploy through hosted APIs like Fireworks or Together AI.

    Tips

    • Use as an open-weight alternative for creative writing when you can avoid Ollama's 4-bit quants and short context defaults, or deploy through hosted APIs like Fireworks or Together AI.
      Source 5
      “Ranks #12 of 146 on LMArena's creative-writing category (Elo 1461), based on blind human preference votes.”
      LMArena creative-writing categoryOpen original ↗
      Source 6
      “Ollama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.”

    Watch out for

    • Avoid Ollama deployments for complex multi-turn roleplay; the default 4-bit quantization and short context window breaks coherence on anything beyond simple chat.
      Source 6
      “Ollama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.”

Frequently asked

What is the top-ranked model for Chat & Roleplay?
Anthropic: Claude Opus 4.6 ranks first in the current evidence-weighted comparison. Use for high-stakes creative writing and roleplay where human judges consistently preferred its output over nearly all competitors.[1]
What is an alternative to Anthropic: Claude Opus 4.6?
OpenAI: GPT-5.6 Sol is the next-ranked option. Consider if you already have OpenAI infrastructure and need creative writing capability, though human judges preferred 15 other models.[2]

Sources

  1. 1

    “Ranks #2 of 146 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.”

    LMArena creative-writing category · Benchmark · Sep 13, 2026
  2. 2

    “Ranks #16 of 146 on LMArena's creative-writing category (Elo 1452), based on blind human preference votes.”

    LMArena creative-writing category · Benchmark · Sep 13, 2026
  3. 3

    “I created a connection profile that used Openrouter as a provider to use Gemini 3.5 flash lite for summarization, and then I would switch to my main roleplay connection profile which used Claude 4.6. It seems that switch disconnects the profile from Openrouter. So, when summarization activated, it would fail.”

    Cheesedozer · GitHub · Aug 2, 2026
  4. 4

    “Ranks #22 of 146 on LMArena's creative-writing category (Elo 1448), based on blind human preference votes.”

    LMArena creative-writing category · Benchmark · Sep 13, 2026
  5. 5

    “Ranks #12 of 146 on LMArena's creative-writing category (Elo 1461), based on blind human preference votes.”

    LMArena creative-writing category · Benchmark · Sep 13, 2026
  6. 6

    “Ollama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.”

    kgeist · Hacker News · Jul 13, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.