Recommendation for OpenHands

OpenHands

Our top recommendation for OpenHands, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use for repository-level bug fixing where SWE-rebench's continuously refreshed issues match your codebase complexity. OpenAI: GPT-5.6 Luna is the next-ranked alternative. Its currently supported evidence is cautionary: Expect poor human preference on multi-step tasks, ranking 18th with a negative score on LMArena's agentic arena.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
3
Revision
v69

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

6

task-weighted

Winner coverage

79%

intended feed weight

Largest provider share

1 of 2

Anthropic

Provisional source breadth. 2 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 50%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
Berkeley Function Calling
32%
not measured3/20
LMArena Agent
23%
#217/20
price weight
15%
93/10020/20
Terminal-Bench 2.1
15%
#1212/20
SWE-rebench
10%
#116/20
OpenRouter usage
5%
79/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic50%
  • Anthropic1 model
  • OpenAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
72
79%no linked practitioner threads#1 SWE-rebench · #2 LMArena Agent
02GPT-5.6 LunaOpenAI
67
79%no linked practitioner threads#4 Terminal-Bench 2.1 · #20 LMArena Agent

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Leads on SWE-rebench with 64.5% resolution rate and ranks second in human preference for agentic tool use.

    Best when: Use for repository-level bug fixing where SWE-rebench's continuously refreshed issues match your codebase complexity.

    Tips

    • Use for repository-level bug fixing where SWE-rebench's continuously refreshed issues match your codebase complexity.
      Source 1
      “Resolves 64.5045045045045% ± 1.4130078505728048 on SWE-rebench (#1 of 13) using tools, a continuously refreshed repository-issue evaluation with configuration recorded separately from the model.”
    • Deploy when human-rated multi-step task performance matters, as it scores 8.8 on LMArena's agentic arena.
      Source 3
      “Ranks #2 of 36 on LMArena's agentic arena (score 8.8), measuring tool-use and multi-step task performance from human preference.”
      LMArena agentic arenaOpen original ↗
  2. Scores below median on both agentic preference and SWE-rebench, with a negative LMArena score and 43.6% issue resolution.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Expect poor human preference on multi-step tasks, ranking 18th with a negative score on LMArena's agentic arena.
      Source 2
      “Ranks #18 of 36 on LMArena's agentic arena (score -0.4), measuring tool-use and multi-step task performance from human preference.”
      LMArena agentic arenaOpen original ↗

Frequently asked

What is the top-ranked model for OpenHands?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for repository-level bug fixing where SWE-rebench's continuously refreshed issues match your codebase complexity.[1]

Sources

  1. 1

    “Resolves 64.5045045045045% ± 1.4130078505728048 on SWE-rebench (#1 of 13) using tools, a continuously refreshed repository-issue evaluation with configuration recorded separately from the model.”

    SWE-rebench · Benchmark · Jul 1, 2026
  2. 2

    “Ranks #18 of 36 on LMArena's agentic arena (score -0.4), measuring tool-use and multi-step task performance from human preference.”

    LMArena agentic arena · Benchmark · Sep 15, 2026
  3. 3

    “Ranks #2 of 36 on LMArena's agentic arena (score 8.8), measuring tool-use and multi-step task performance from human preference.”

    LMArena agentic arena · Benchmark · Sep 15, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.