Recommendation for Marketing & copy
Marketing Copy
Our top recommendation for Marketing Copy, based on the public evidence we track, is Anthropic: Claude Fable 5.[1] Use for high-stakes conversion copy where blind human preference predicts real-world engagement, given its #1 ranking in head-to-head voting. Anthropic: Claude Opus 4.6 is the next-ranked alternative.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 10
- Revision
- v78
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
2 of 6
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LiveBench Instruction Following | 30% | #5 | 20/21 |
| LMArena Instruction Following | 25% | #4 | 21/21 |
| LMArena Text | 20% | #1 | 21/21 |
| LMArena Creative Writing | 15% | #4 | 21/21 |
| OpenRouter usage | 10% | 79/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic2 models
- Google1 model
- Meta1 model
- OpenAI1 model
- Qwen1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 84 | 100% | no linked practitioner threads | #1 LMArena Text · #4 LMArena Creative Writing |
| 02 | Claude Opus 4.6Anthropic | 80 | 100% | no linked practitioner threads | #1 LMArena Instruction Following · #2 LMArena Creative Writing |
| 03 | GPT-5.6 SolOpenAI | 80 | 100% | no linked practitioner threads | #12 LMArena Instruction Following · #13 LMArena Text |
| 04 | Gemini 3.6 FlashGoogle | 80 | 100% | no linked practitioner threads | #8 LiveBench Instruction Following · #10 LMArena Creative Writing |
| 05 | Muse Spark 1.2Meta | 77 | 100% | no linked practitioner threads | #4 LMArena Text · #10 LiveBench Instruction Following |
| 06 | Qwen3.7 MaxQwen | 76 | 100% | no linked practitioner threads | #11 LiveBench Instruction Following · #19 LMArena Creative Writing |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Leads the LMArena text arena with the highest human preference score, suggesting strong alignment with how people judge persuasive, natural-sounding copy.
Best when: Use for high-stakes conversion copy where blind human preference predicts real-world engagement, given its #1 ranking in head-to-head voting.
Tips
- Use for high-stakes conversion copy where blind human preference predicts real-world engagement, given its #1 ranking in head-to-head voting.
- Deploy for instruction-following tasks like paraphrasing and simplifying existing marketing materials, where it scores 75.77% on LiveBench.
Anthropic: Claude Opus 4.6 ranks #2 of 146 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.
Best when: Consider only after reviewing the cited caution.
Shows a 13-rank gap between LMArena preference (#13) and LiveBench instruction-following (#18), indicating competent but not standout copy generation.
Best when: Use when you need paraphrasing and simplification at 71.85% accuracy on LiveBench, adequate for routine product description rewrites.
Tips
- Use when you need paraphrasing and simplification at 71.85% accuracy on LiveBench, adequate for routine product description rewrites.
Watch out for
- Expect lower human preference scores than top-tier models, with Elo 1483 placing it outside the preference elite for persuasive content.
Matches top models on LiveBench instruction-following (75.37%) while sitting mid-pack in LMArena preference, suggesting capable but less charismatic output.
Best when: Use for structured copy tasks like summarizing product specs into descriptions, where its #8 instruction-following score indicates reliable adherence to format constraints.
Tips
- Use for structured copy tasks like summarizing product specs into descriptions, where its #8 instruction-following score indicates reliable adherence to format constraints.
Watch out for
- Expect lower blind human preference than top-ranked models (Elo 1480 vs 1506), which may translate to less engaging ad hook performance.
Ranks #4 in LMArena human preference with strong instruction-following scores, placing it in the top tier for copy that follows creative briefs.
Best when: Use for iterative copy refinement where LiveBench's paraphrasing and simplification tasks predict workflow fit, scoring 74.33%.
Tips
- Use for iterative copy refinement where LiveBench's paraphrasing and simplification tasks predict workflow fit, scoring 74.33%.
- Deploy when you need near-top-tier human preference (Elo 1500) with potentially different cost or latency profiles than Anthropic models.
Ranks #23 in LMArena preference with mid-tier instruction-following, suggesting functional copy generation without standout human appeal.
Best when: Use for bulk copy generation where LiveBench's 74.04% instruction-following score provides adequate control over output structure.
Tips
- Use for bulk copy generation where LiveBench's 74.04% instruction-following score provides adequate control over output structure.
Watch out for
- Expect significantly lower blind human preference than top models, with 33-point Elo gap indicating less naturally engaging prose.
Frequently asked
- What is the top-ranked model for Marketing Copy?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for high-stakes conversion copy where blind human preference predicts real-world engagement, given its #1 ranking in head-to-head voting.[1]
Sources
- 1
“Ranks #1 of 146 on LMArena's overall text arena (Elo 1506), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 2
“Scores 75.77% on LiveBench Instruction Following (#5 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 3
“Scores 71.85% on LiveBench Instruction Following (#18 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 4
“Ranks #13 of 146 on LMArena's overall text arena (Elo 1483), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 5
“Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 6
“Ranks #17 of 146 on LMArena's overall text arena (Elo 1480), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 7
“Scores 74.33% on LiveBench Instruction Following (#10 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 8
“Ranks #4 of 146 on LMArena's overall text arena (Elo 1500), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 9
“Scores 74.04% on LiveBench Instruction Following (#12 of 58), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 10
“Ranks #23 of 146 on LMArena's overall text arena (Elo 1473), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.