Recommendation for Everyday translation
Everyday Translation
Our top recommendation for Everyday Translation, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use when you need translations that human raters prefer over 145 other models, as it holds the top Elo score on LMArena's blind text arena. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Choose for translation workflows where LiveBench Language scores guide selection, as its 87.68% places 6th of 58 tested models.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 7
- Revision
- v78
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
44%
intended feed weight
Largest provider share
2 of 5
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| WMT translationunavailable | 45% | feed unavailable | 0/20 |
| LiveBench Language | 20% | #1 | 18/20 |
| LMArena Text | 15% | #1 | 19/20 |
| LiveBench Instruction Following | 10% | #5 | 18/20 |
| OpenRouter usage | 10% | 79/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic2 models
- OpenAI1 model
- xAI1 model
- xiaomi1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 57 | 44% | no linked practitioner threads | #1 LiveBench Language · #1 LMArena Text |
| 02 | GPT-5.6 SolOpenAI | 57 | 44% | no linked practitioner threads | #6 LiveBench Language · #13 LMArena Text |
| 03 | MiMo-V2.6-Proxiaomi | 52 | 12% | 1 threads · 1 families · 0 cautions | OpenRouter usage 91/100 normalized |
| 04 | Claude Opus 4.6Anthropic | 52 | 44% | no linked practitioner threads | #2 LMArena Text · #15 LiveBench Language |
| 05 | Grok 4.6xAI | 52 | 44% | no linked practitioner threads | #14 LiveBench Language · #16 LiveBench Instruction Following |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Leads on both human preference and language manipulation benchmarks, suggesting strong naturalness and accuracy for everyday translation tasks.
Best when: Use when you need translations that human raters prefer over 145 other models, as it holds the top Elo score on LMArena's blind text arena.
Tips
- Use when you need translations that human raters prefer over 145 other models, as it holds the top Elo score on LMArena's blind text arena.
- Deploy for language manipulation tasks where benchmark scores matter, given its #1 ranking on LiveBench Language at 90.68%.
Solid benchmark performer with strong language manipulation scores, though human preference rankings place it mid-pack among top contenders.
Best when: Choose for translation workflows where LiveBench Language scores guide selection, as its 87.68% places 6th of 58 tested models.
Tips
- Choose for translation workflows where LiveBench Language scores guide selection, as its 87.68% places 6th of 58 tested models.
Open-weight option with explicit community demand for translation integration and competitive pricing.
Best when: Select for cost-sensitive high-volume translation, with Pro tier at ¥3/¥6 per million tokens and Flash at ¥1/¥2 as of September 2026.
Tips
- Select for cost-sensitive high-volume translation, with Pro tier at ¥3/¥6 per million tokens and Flash at ¥1/¥2 as of September 2026.
Watch out for
- Handle thinking-mode output carefully, as the model defaults to chain-of-thought and requires explicit configuration to extract clean translations.
Nearly matches the top human preference score while trading some language manipulation accuracy for that naturalness.
Best when: Select when human-like output quality matters most, as its Elo 1505 ranks #2 on LMArena's blind preference voting.
Tips
- Select when human-like output quality matters most, as its Elo 1505 ranks #2 on LMArena's blind preference voting.
Watch out for
- Watch for lower performance on structured language tasks, with LiveBench Language at 83.27% ranking 17th of 58.
Mid-tier language manipulation performance without human preference data to confirm everyday translation quality.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Verify output quality manually, as its 16th-place LiveBench Language score (83.7%) lacks supporting human preference evidence.
Frequently asked
- What is the top-ranked model for Everyday Translation?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use when you need translations that human raters prefer over 145 other models, as it holds the top Elo score on LMArena's blind text arena.[1]
- What is an alternative to Anthropic: Claude Fable 5?
- OpenAI: GPT-5.6 Sol is the next-ranked option. Choose for translation workflows where LiveBench Language scores guide selection, as its 87.68% places 6th of 58 tested models.[2]
Sources
- 1
“Ranks #1 of 146 on LMArena's overall text arena (Elo 1506), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 2
“Scores 87.68% on LiveBench Language (#6 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 3
“Scores 90.68% on LiveBench Language (#1 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 4
“### 您希望的更新和改进是什么 | Update or Improve 希望在翻译服务中提供小米 MiMo 官方 API 的配置入口,支持选择 `mimo-v2.6-flash` 和 `mimo-v2.6-pro`。用户填写自己的 API Key 后,可以将这两个模型用于网页、PDF 等内容翻译。 MiMo-V2.6 系列已具备值得在翻译场景与现有常用大模型比较的能力,按量计费也适合频繁翻译。以 2026 年 9 月 23 日的官方价格为例,每百万非缓存输入 / 输出 tokens:Flash 为 ¥1 / ¥2,Pro 为 ¥3 / ¥6。希望用户可以直接选用这两个模型,按自己的译文质量和费用需求做选择。 建议接入小米官方的 OpenAI 兼容接口: https://api.xiaomimimo.com/v1/chat/completions MiMo-V2.6 默认启用 thinking;翻译场景希望能正确提取译文,并提供关闭 thinking 的配置。 ### 补充说明 | Additional context 此前 #3596 讨论过接入小米 MiMo 平台;#3947 讨…”
Gnaix-39 · GitHub · Sep 23, 2026 - 5
“Ranks #2 of 146 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 13, 2026 - 6
“Scores 83.27% on LiveBench Language (#17 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 7
“Scores 83.7% on LiveBench Language (#16 of 58), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.