Recommendation for Translation

Translation

Our top recommendation for Translation, based on the public evidence we track, is OpenAI: GPT-6 Astra.[1][2][3] Deploy for translation pipelines where automated evaluation of language structure matters, given its #3 ranking on LiveBench Language. Watch out: Expect lower end-user satisfaction in consumer-facing applications, as human voters ranked it #18 in blind preference. Anthropic: Claude Fable 5 is the next-ranked alternative. Use when you need translations that human evaluators prefer over competitors, as it ranks #1 in blind preference votes on LMArena.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
11
Revision
v80

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

22

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

44%

intended feed weight

Largest provider share

3 of 6

Anthropic

Provisional source breadth. 2 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 55%.

Sources evaluated

The task sets these weights before any model is scored.

winner: GPT-6 Astra
Evaluation feedWeightWinner resultField measured
WMT translationunavailable
45%
feed unavailable0/22
LiveBench Language
20%
#318/22
LMArena Text
15%
#1622/22
LiveBench Instruction Following
10%
#718/22
OpenRouter usage
10%
97/10022/22

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic50%
  • Anthropic3 models
  • OpenAI2 models
  • Google1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01GPT-6 AstraOpenAI
58
44%no linked practitioner threads#3 LiveBench Language · #7 LiveBench Instruction Following
02Claude Fable 5Anthropic
57
44%no linked practitioner threads#1 LiveBench Language · #1 LMArena Text
03GPT-5.6 SolOpenAI
57
44%no linked practitioner threads#6 LiveBench Language · #13 LMArena Text
04Gemini 3.6 FlashGoogle
54
44%no linked practitioner threads#8 LiveBench Instruction Following · #13 LiveBench Language
05Claude Opus 4.6Anthropic
52
44%no linked practitioner threads#2 LMArena Text · #15 LiveBench Language
06Claude Opus 4.8Anthropic
52
44%no linked practitioner threads#15 LiveBench Instruction Following · #15 LMArena Text

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Shows strong objective language manipulation performance relative to its lower human preference ranking, suggesting technically capable translations that may read less naturally.

    Best when: Deploy for translation pipelines where automated evaluation of language structure matters, given its #3 ranking on LiveBench Language.

    Tips

    • Deploy for translation pipelines where automated evaluation of language structure matters, given its #3 ranking on LiveBench Language.
      Source 1
      “Scores 89.43% on LiveBench Language (#3 of 58), an objective evaluation of language manipulation tasks.”
      LiveBench LanguageOpen original ↗

    Watch out for

    • Expect lower end-user satisfaction in consumer-facing applications, as human voters ranked it #18 in blind preference.
      Source 2
      “Ranks #18 of 146 on LMArena's overall text arena (Elo 1480), based on blind human preference votes.”
      LMArena text arenaOpen original ↗
  2. Leads both human preference rankings and objective language manipulation benchmarks, suggesting strong translation fluency and multilingual output quality.

    Best when: Use when you need translations that human evaluators prefer over competitors, as it ranks #1 in blind preference votes on LMArena.

    Tips

    • Use when you need translations that human evaluators prefer over competitors, as it ranks #1 in blind preference votes on LMArena.
      Source 3
      “Ranks #1 of 146 on LMArena's overall text arena (Elo 1506), based on blind human preference votes.”
      LMArena text arenaOpen original ↗
    • Deploy for structured language tasks where objective scoring matters, given its top position on LiveBench Language.
      Source 4
      “Scores 90.68% on LiveBench Language (#1 of 58), an objective evaluation of language manipulation tasks.”
      LiveBench LanguageOpen original ↗
  3. Holds a mid-tier position in both preference and objective benchmarks, indicating competent but not leading translation capabilities.

    Best when: Use as a fallback when top-ranked models are unavailable, with #13 preference ranking and #6 language manipulation score providing reasonable confidence.

    Tips

    • Use as a fallback when top-ranked models are unavailable, with #13 preference ranking and #6 language manipulation score providing reasonable confidence.
      Source 5
      “Ranks #13 of 146 on LMArena's overall text arena (Elo 1483), based on blind human preference votes.”
      LMArena text arenaOpen original ↗
      Source 6
      “Scores 87.68% on LiveBench Language (#6 of 58), an objective evaluation of language manipulation tasks.”
      LiveBench LanguageOpen original ↗
  4. Sits in the middle of the pack on both preference and objective metrics without distinguishing strengths for translation tasks.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Look elsewhere for high-stakes translation work, as its #15 LiveBench Language score and #17 preference ranking indicate no particular advantage.
      Source 7
      “Ranks #17 of 146 on LMArena's overall text arena (Elo 1480), based on blind human preference votes.”
      LMArena text arenaOpen original ↗
      Source 8
      “Scores 83.9% on LiveBench Language (#15 of 58), an objective evaluation of language manipulation tasks.”
      LiveBench LanguageOpen original ↗
  5. Matches near-top human preference scores but shows weaker objective language manipulation performance, creating a trade-off between perceived fluency and measured accuracy.

    Best when: Choose when human-like translation tone is prioritized, as it ranks #2 in blind preference voting.

    Tips

    • Choose when human-like translation tone is prioritized, as it ranks #2 in blind preference voting.
      Source 9
      “Ranks #2 of 146 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.”
      LMArena text arenaOpen original ↗

    Watch out for

    • Avoid for precision-critical translation workflows where benchmarked language manipulation accuracy matters, as it scores only #17 on LiveBench Language.
      Source 10
      “Scores 83.27% on LiveBench Language (#17 of 58), an objective evaluation of language manipulation tasks.”
      LiveBench LanguageOpen original ↗
  6. Lacks objective benchmark evidence for language tasks, with only moderate human preference data available.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Treat translation quality as unverified, since no LiveBench Language scores are provided to assess actual multilingual capability.
      Source 11
      “Ranks #16 of 146 on LMArena's overall text arena (Elo 1481), based on blind human preference votes.”
      LMArena text arenaOpen original ↗

Frequently asked

What is the top-ranked model for Translation?
OpenAI: GPT-6 Astra ranks first in the current evidence-weighted comparison. Deploy for translation pipelines where automated evaluation of language structure matters, given its #3 ranking on LiveBench Language.[1]
What should I watch out for with OpenAI: GPT-6 Astra?
Expect lower end-user satisfaction in consumer-facing applications, as human voters ranked it #18 in blind preference.[2]
What is an alternative to OpenAI: GPT-6 Astra?
Anthropic: Claude Fable 5 is the next-ranked option. Use when you need translations that human evaluators prefer over competitors, as it ranks #1 in blind preference votes on LMArena.[3]

Sources

  1. 1

    “Scores 89.43% on LiveBench Language (#3 of 58), an objective evaluation of language manipulation tasks.”

    LiveBench Language · Benchmark · Jun 25, 2026
  2. 2

    “Ranks #18 of 146 on LMArena's overall text arena (Elo 1480), based on blind human preference votes.”

    LMArena text arena · Benchmark · Sep 13, 2026
  3. 3

    “Ranks #1 of 146 on LMArena's overall text arena (Elo 1506), based on blind human preference votes.”

    LMArena text arena · Benchmark · Sep 13, 2026
  4. 4

    “Scores 90.68% on LiveBench Language (#1 of 58), an objective evaluation of language manipulation tasks.”

    LiveBench Language · Benchmark · Jun 25, 2026
  5. 5

    “Ranks #13 of 146 on LMArena's overall text arena (Elo 1483), based on blind human preference votes.”

    LMArena text arena · Benchmark · Sep 13, 2026
  6. 6

    “Scores 87.68% on LiveBench Language (#6 of 58), an objective evaluation of language manipulation tasks.”

    LiveBench Language · Benchmark · Jun 25, 2026
  7. 7

    “Ranks #17 of 146 on LMArena's overall text arena (Elo 1480), based on blind human preference votes.”

    LMArena text arena · Benchmark · Sep 13, 2026
  8. 8

    “Scores 83.9% on LiveBench Language (#15 of 58), an objective evaluation of language manipulation tasks.”

    LiveBench Language · Benchmark · Jun 25, 2026
  9. 9

    “Ranks #2 of 146 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.”

    LMArena text arena · Benchmark · Sep 13, 2026
  10. 10

    “Scores 83.27% on LiveBench Language (#17 of 58), an objective evaluation of language manipulation tasks.”

    LiveBench Language · Benchmark · Jun 25, 2026
  11. 11

    “Ranks #16 of 146 on LMArena's overall text arena (Elo 1481), based on blind human preference votes.”

    LMArena text arena · Benchmark · Sep 13, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.