Recommendation for Goose

Goose

Our top recommendation for Goose, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use for repository automation workflows where resolving real GitHub issues matters: it tops SWE-rebench at 64.5% resolution rate, the highest among all candidates. OpenAI: GPT-5.6 Luna is the next-ranked alternative. Its currently supported evidence is cautionary: Expect poor human-rated tool-use quality: it ranks #18 on LMArena's agentic arena with a negative score of -0.4, indicating below-average performance.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
3
Revision
v69

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

6

task-weighted

Winner coverage

79%

intended feed weight

Largest provider share

1 of 2

Anthropic

Provisional source breadth. 2 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 50%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
Berkeley Function Calling
32%
not measured3/20
LMArena Agent
23%
#217/20
price weight
15%
93/10020/20
Terminal-Bench 2.1
15%
#1212/20
SWE-rebench
10%
#116/20
OpenRouter usage
5%
79/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic50%
  • Anthropic1 model
  • OpenAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
72
79%no linked practitioner threads#1 SWE-rebench · #2 LMArena Agent
02GPT-5.6 LunaOpenAI
67
79%no linked practitioner threads#4 Terminal-Bench 2.1 · #20 LMArena Agent

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 leads on both human preference rankings and automated repository issue resolution, making it the strongest candidate for multi-step tool use in Goose.

    Best when: Use for repository automation workflows where resolving real GitHub issues matters: it tops SWE-rebench at 64.5% resolution rate, the highest among all candidates.

    Tips

    • Use for repository automation workflows where resolving real GitHub issues matters: it tops SWE-rebench at 64.5% resolution rate, the highest among all candidates.
      Source 1
      “Resolves 64.5045045045045% ± 1.4130078505728048 on SWE-rebench (#1 of 13) using tools, a continuously refreshed repository-issue evaluation with configuration recorded separately from the model.”
    • Deploy when human-rated tool-use quality is the priority: it ranks #2 on LMArena's agentic arena with a score of 8.8, indicating strong multi-step task performance.
      Source 3
      “Ranks #2 of 36 on LMArena's agentic arena (score 8.8), measuring tool-use and multi-step task performance from human preference.”
      LMArena agentic arenaOpen original ↗
  2. GPT-5.6 Luna underperforms on both agentic benchmarks, with negative human preference scores and sub-50% repository issue resolution.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Expect poor human-rated tool-use quality: it ranks #18 on LMArena's agentic arena with a negative score of -0.4, indicating below-average performance.
      Source 2
      “Ranks #18 of 36 on LMArena's agentic arena (score -0.4), measuring tool-use and multi-step task performance from human preference.”
      LMArena agentic arenaOpen original ↗

Frequently asked

What is the top-ranked model for Goose?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for repository automation workflows where resolving real GitHub issues matters: it tops SWE-rebench at 64.5% resolution rate, the highest among all candidates.[1]

Sources

  1. 1

    “Resolves 64.5045045045045% ± 1.4130078505728048 on SWE-rebench (#1 of 13) using tools, a continuously refreshed repository-issue evaluation with configuration recorded separately from the model.”

    SWE-rebench · Benchmark · Jul 1, 2026
  2. 2

    “Ranks #18 of 36 on LMArena's agentic arena (score -0.4), measuring tool-use and multi-step task performance from human preference.”

    LMArena agentic arena · Benchmark · Sep 15, 2026
  3. 3

    “Ranks #2 of 36 on LMArena's agentic arena (score 8.8), measuring tool-use and multi-step task performance from human preference.”

    LMArena agentic arena · Benchmark · Sep 15, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.