Recommendation for Literary & nuanced
Literary Translation
Our top recommendation for Literary Translation, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use when translating poetry or classical texts where meter and register fidelity matter, as it identified and continued William Cowper's 18th-century translation of Homer from just two lines without web search. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Acceptable fallback for general translation workflows where LiveBench Language scores above 85% indicate competent handling of language tasks.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 11
- Revision
- v76
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
44%
intended feed weight
Largest provider share
3 of 7
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| WMT translationunavailable | 45% | feed unavailable | 0/21 |
| LiveBench Language | 20% | #1 | 19/21 |
| LMArena Text | 15% | #1 | 21/21 |
| LiveBench Instruction Following | 10% | #5 | 19/21 |
| OpenRouter usage | 10% | 79/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic3 models
- deepseek1 model
- Google1 model
- OpenAI1 model
- xAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 57 | 44% | 1 threads · 1 families · 0 cautions | #1 LiveBench Language · #1 LMArena Text |
| 02 | GPT-5.6 SolOpenAI | 57 | 44% | no linked practitioner threads | #6 LiveBench Language · #13 LMArena Text |
| 03 | Gemini 3.6 FlashGoogle | 54 | 44% | no linked practitioner threads | #8 LiveBench Instruction Following · #13 LiveBench Language |
| 04 | DeepSeek V4 Flash 0423deepseek | 52 | 44% | 1 threads · 1 families · 0 cautions | #40 LiveBench Instruction Following · #50 LiveBench Language |
| 05 | Claude Opus 4.6Anthropic | 52 | 44% | no linked practitioner threads | #2 LMArena Text · #15 LiveBench Language |
| 06 | Claude Opus 4.8Anthropic | 52 | 44% | no linked practitioner threads | #15 LiveBench Instruction Following · #15 LMArena Text |
| 07 | Grok 4.6xAI | 52 | 44% | no linked practitioner threads | #14 LiveBench Language · #16 LiveBench Instruction Following |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Leads on both human preference and language manipulation benchmarks, with demonstrated ability to continue classical English verse from minimal context.
Best when: Use when translating poetry or classical texts where meter and register fidelity matter, as it identified and continued William Cowper's 18th-century translation of Homer from just two lines without web search.
Tips
- Use when translating poetry or classical texts where meter and register fidelity matter, as it identified and continued William Cowper's 18th-century translation of Homer from just two lines without web search.
- Deploy for high-stakes literary work where benchmark reliability predicts quality: it scores 90.68% on LiveBench Language, ranking first among 58 models on objective language manipulation tasks.
Mid-tier on human preference with solid but not leading language manipulation scores, lacking specific literary translation evidence.
Best when: Acceptable fallback for general translation workflows where LiveBench Language scores above 85% indicate competent handling of language tasks.
Tips
- Acceptable fallback for general translation workflows where LiveBench Language scores above 85% indicate competent handling of language tasks.
Watch out for
- Expect less refined output than top-ranked models, as its #13 LMArena position and 87.68% language score place it outside the top tier on both metrics.
Lower-tier on both human preference and language manipulation benchmarks with no literary-specific capabilities demonstrated.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Avoid for nuanced literary translation, as its 83.9% LiveBench Language score and #17 LMArena rank place it below multiple competitors on both metrics.
Open-weight candidate with only infrastructure bug reports available, no translation quality or benchmark evidence.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Expect formatting instability in streaming workflows, as reported whitespace and markdown corruption affects output structure.
- No quality validation exists, as evidence covers only routing infrastructure bugs with no benchmark, human preference, or translation task data.
Ranks second on human preference with strong overall Elo, though its language manipulation score trails Fable by seven percentage points.
Best when: Consider for prose translation where human judges consistently preferred output, given its #2 ranking on LMArena's blind text arena.
Tips
- Consider for prose translation where human judges consistently preferred output, given its #2 ranking on LMArena's blind text arena.
Watch out for
- Verify literary nuance manually, as its 83.27% on LiveBench Language places 17th, suggesting weaker structured language manipulation than top competitors.
Human preference data only, with no benchmark scores or literary-specific evidence available.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Lack objective validation for language manipulation, as no LiveBench or literary task evidence accompanies its mid-table LMArena ranking.
Benchmarked on language manipulation but lacks human preference ranking or literary translation examples.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Treat with caution for literary work, as its 83.7% LiveBench Language score ranks 16th and no human preference or verse-continuation evidence exists.
Frequently asked
- What is the top-ranked model for Literary Translation?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use when translating poetry or classical texts where meter and register fidelity matter, as it identified and continued William Cowper's 18th-century translation of Homer from just two lines without web search.[1]
- What is an alternative to Anthropic: Claude Fable 5?
- OpenAI: GPT-5.6 Sol is the next-ranked option. Acceptable fallback for general translation workflows where LiveBench Language scores above 85% indicate competent handling of language tasks.[2]
Sources
- 1
“> The only thing Fable was given was a clean of copy of the ancient Greek. This is not a bastardization of other translations FWIW I gave GPT-5.5 two lines of William Cowper's translation (long out of copyright to avoid any filters) without web search and not only was it able to identify it, it was able to continue it verbatim for roughly 30 lines at which point it went off course, but only by skipping like two books ahead. If I had supplied the original text as an anchor it presumably could ha…”
magicalist · Hacker News · Aug 7, 2026 - 2
“Scores 87.68% on LiveBench Language (#6 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 3
“Scores 90.68% on LiveBench Language (#1 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 4
“Ranks #13 of 146 on LMArena's overall text arena (Elo 1483), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 5
“Ranks #17 of 146 on LMArena's overall text arena (Elo 1480), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 6
“Scores 83.9% on LiveBench Language (#15 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 7
“### OmniRoute Version 3.8.49 (docker image `diegosouzapw/omniroute:latest`) ### Installation Method docker ### Operating System Linux (Proxmox LXC) ### Provider(s) Involved opencode-go, opencode-zen, openrouter (any streaming provider) ### Model(s) Involved deepseek-v4-flash, ox-alpha-free (multiple models — not model-specific) ### Client Tool Claude Code CLI (Anthropic `/v1/messages` streaming) ### Description Streaming responses through OmniRoute get their whitespace/markdown structure corrup…”
vinnyduke · GitHub · Aug 26, 2026 - 8
“Ranks #2 of 146 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 9
“Scores 83.27% on LiveBench Language (#17 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 10
“Ranks #16 of 146 on LMArena's overall text arena (Elo 1481), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 11
“Scores 83.7% on LiveBench Language (#16 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.