Recommendation for OpenHands
OpenHands
Our top recommendation for OpenHands, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use for repository-level bug fixing where SWE-rebench's continuously refreshed issues match your codebase complexity. OpenAI: GPT-5.6 Luna is the next-ranked alternative. Its currently supported evidence is cautionary: Expect poor human preference on multi-step tasks, ranking 18th with a negative score on LMArena's agentic arena.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 3
- Revision
- v69
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
6
task-weighted
Winner coverage
79%
intended feed weight
Largest provider share
1 of 2
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| Berkeley Function Calling | 32% | not measured | 3/20 |
| LMArena Agent | 23% | #2 | 17/20 |
| price weight | 15% | 93/100 | 20/20 |
| Terminal-Bench 2.1 | 15% | #12 | 12/20 |
| SWE-rebench | 10% | #1 | 16/20 |
| OpenRouter usage | 5% | 79/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic1 model
- OpenAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 72 | 79% | no linked practitioner threads | #1 SWE-rebench · #2 LMArena Agent |
| 02 | GPT-5.6 LunaOpenAI | 67 | 79% | no linked practitioner threads | #4 Terminal-Bench 2.1 · #20 LMArena Agent |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Leads on SWE-rebench with 64.5% resolution rate and ranks second in human preference for agentic tool use.
Best when: Use for repository-level bug fixing where SWE-rebench's continuously refreshed issues match your codebase complexity.
Tips
- Use for repository-level bug fixing where SWE-rebench's continuously refreshed issues match your codebase complexity.
- Deploy when human-rated multi-step task performance matters, as it scores 8.8 on LMArena's agentic arena.
Scores below median on both agentic preference and SWE-rebench, with a negative LMArena score and 43.6% issue resolution.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Expect poor human preference on multi-step tasks, ranking 18th with a negative score on LMArena's agentic arena.
Frequently asked
- What is the top-ranked model for OpenHands?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for repository-level bug fixing where SWE-rebench's continuously refreshed issues match your codebase complexity.[1]
Sources
- 1
“Resolves 64.5045045045045% ± 1.4130078505728048 on SWE-rebench (#1 of 13) using tools, a continuously refreshed repository-issue evaluation with configuration recorded separately from the model.”
SWE-rebench · Benchmark · Jul 1, 2026 - 2
“Ranks #18 of 36 on LMArena's agentic arena (score -0.4), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Sep 15, 2026 - 3
“Ranks #2 of 36 on LMArena's agentic arena (score 8.8), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Sep 15, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.