Recommendation for Competition math

Competition Math

Our top recommendation for Competition Math, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use when you need a model with validated competition math performance on both human preference and objective benchmarks, as it scores 95.99% on LiveBench Mathematics. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Deploy for competition math where LiveBench scores matter, as its 96.2% places it in the top tier of objectively evaluated models.

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
8
Revision
v78

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

3 of 6

Anthropic

Provisional source breadth. 5 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 33%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
LiveBench Mathematics
35%
#621/21
LiveBench Reasoning
25%
#1021/21
LMArena Math
25%
#319/21
LMArena Instruction Following
10%
#420/21
OpenRouter usage
5%
79/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic50%
  • Anthropic3 models
  • OpenAI2 models
  • Qwen1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
85
100%no linked practitioner threads#3 LMArena Math · #4 LMArena Instruction Following
02GPT-5.6 SolOpenAI
82
100%1 threads · 1 families · 1 cautions#4 LiveBench Reasoning · #5 LiveBench Mathematics
03Claude Opus 4.6Anthropic
79
100%no linked practitioner threads#1 LMArena Instruction Following · #5 LMArena Math
04GPT-5.6 LunaOpenAI
78
100%1 threads · 1 families · 0 cautions#25 LiveBench Reasoning · #34 LiveBench Mathematics
05Qwen3.8 27BQwen
75
100%1 threads · 1 families · 0 cautions#25 LMArena Math · #37 LMArena Instruction Following
06Claude Opus 5.5Anthropic
63
44%no linked practitioner threads#1 LiveBench Mathematics · #2 LiveBench Reasoning

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 places 4th in human preference rankings for math and 6th on objectively scored competition tasks, showing strong but not top-tier performance on rigorous olympiad problems.

    Best when: Use when you need a model with validated competition math performance on both human preference and objective benchmarks, as it scores 95.99% on LiveBench Mathematics.

    Tips

    • Use when you need a model with validated competition math performance on both human preference and objective benchmarks, as it scores 95.99% on LiveBench Mathematics.
      Source 1
      “Scores 95.99% on LiveBench Mathematics (#6 of 58), using objectively scored competition and olympiad-style tasks.”
      LiveBench MathematicsOpen original ↗
  2. GPT-5.6 Sol achieves 96.2% on LiveBench Mathematics and ranks 5th, though an independent numerical review identified a floating-point underflow bug in its optimization routines that could affect certain mathematical derivations.

    Best when: Deploy for competition math where LiveBench scores matter, as its 96.2% places it in the top tier of objectively evaluated models.

    Tips

    • Deploy for competition math where LiveBench scores matter, as its 96.2% places it in the top tier of objectively evaluated models.
      Source 2
      “Scores 96.2% on LiveBench Mathematics (#5 of 58), using objectively scored competition and olympiad-style tasks.”
      LiveBench MathematicsOpen original ↗

    Watch out for

    • Watch for numerical stability issues in optimization-heavy derivations, as float32 underflow has been reproduced in Armijo condition calculations with small gradient terms.
      Source 3
      “INDEPENDENT NUMERICAL REVIEW · Z-Switch-Q9V3 / GPT-5.6 Sol · issue #15960 Fresh review found a blocker in the earlier proposed "scale before reduce" repair: it fixes the original overflow witness but can underflow small Armijo terms to zero and accept an insufficiently improving step. Reproduced float32 witness: 200 coordinates, g_i=1e-35, d_i=-7, armijo=1e-4, step=1. The mathematically required Armijo decrease is nonzero, but multiplying g by armijo first flushes the per-coordinate terms to ze…”
      woahwhattheheckOpen original ↗
  3. Claude Opus 4.6 ranks 5th in human preference for math prompts but lacks objective benchmark scores, making its performance on rigorously scored competition problems harder to verify.

    Best when: Consider when human-rated math helpfulness is your priority, as its 1516 Elo in LMArena's maths category reflects strong perceived quality in blind evaluations.

    Tips

    • Consider when human-rated math helpfulness is your priority, as its 1516 Elo in LMArena's maths category reflects strong perceived quality in blind evaluations.
      Source 4
      “Ranks #5 of 139 on LMArena's maths category (Elo 1516), based on blind human preference votes for maths prompts.”
      LMArena maths categoryOpen original ↗
  4. GPT-5.6 Luna has demonstrated a complete, verified solve of a competition problem through an end-to-end proxy test, though no benchmark standings are available to compare against other models.

    Best when: Use when you need a model with proven end-to-end solving capability on actual competition problems, as it passed a full slow-mode solve with dual-judge verification in under 5 minutes.

    Tips

    • Use when you need a model with proven end-to-end solving capability on actual competition problems, as it passed a full slow-mode solve with dual-judge verification in under 5 minutes.
      Source 5
      “## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…”

    Watch out for

    • Note the lack of benchmark comparisons, as no LiveBench or LMArena scores are provided to gauge relative performance against competitors.
      Source 5
      “## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…”
  5. Qwen3.8 27B is an open-weight model with strong reported benchmarks including GPQA Diamond 89.2, though it ranks 26th in human preference for math and was not adopted upstream due to unresolved integration questions.

    Best when: Choose for local deployment where Apache 2.0 licensing matters, as this 27B parameter model offers competitive benchmark scores without API dependencies.

    Tips

    • Choose for local deployment where Apache 2.0 licensing matters, as this 27B parameter model offers competitive benchmark scores without API dependencies.
      Source 6
      “**Standing issue — intentionally left open. Revisit periodically; close only if the model is ruled out for good.** ## What this tracks Whether [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) should replace Gemma 4 26B-A4B as the local model, and what would have to change upstream for that to make sense. ## Why it was not adopted on release (2026-08-15) The benchmarks are excellent — GPQA Diamond 89.2, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0, SWE-bench Pro 61.7, Apache 2.0, 262K na…”

    Watch out for

    • Expect lower human-rated math quality than top closed-weight alternatives, as its 1470 Elo in LMArena's maths category places it well outside the top tier.
      Source 7
      “Ranks #26 of 139 on LMArena's maths category (Elo 1470), based on blind human preference votes for maths prompts.”
      LMArena maths categoryOpen original ↗
  6. Claude Opus 5.5 leads all models on LiveBench Mathematics with 97.08%, achieving the top rank on objectively scored competition and olympiad-style tasks.

    Best when: Prioritize for maximum performance on rigorously evaluated competition math, as its 97.08% and #1 ranking on LiveBench Mathematics exceed all other reported scores.

    Tips

    • Prioritize for maximum performance on rigorously evaluated competition math, as its 97.08% and #1 ranking on LiveBench Mathematics exceed all other reported scores.
      Source 8
      “Scores 97.08% on LiveBench Mathematics (#1 of 58), using objectively scored competition and olympiad-style tasks.”
      LiveBench MathematicsOpen original ↗

Frequently asked

What is the top-ranked model for Competition Math?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use when you need a model with validated competition math performance on both human preference and objective benchmarks, as it scores 95.99% on LiveBench Mathematics.[1]
What is an alternative to Anthropic: Claude Fable 5?
OpenAI: GPT-5.6 Sol is the next-ranked option. Deploy for competition math where LiveBench scores matter, as its 96.2% places it in the top tier of objectively evaluated models.[2]

Sources

  1. 1

    “Scores 95.99% on LiveBench Mathematics (#6 of 58), using objectively scored competition and olympiad-style tasks.”

    LiveBench Mathematics · Benchmark · Jun 25, 2026
  2. 2

    “Scores 96.2% on LiveBench Mathematics (#5 of 58), using objectively scored competition and olympiad-style tasks.”

    LiveBench Mathematics · Benchmark · Jun 25, 2026
  3. 3

    “INDEPENDENT NUMERICAL REVIEW · Z-Switch-Q9V3 / GPT-5.6 Sol · issue #15960 Fresh review found a blocker in the earlier proposed "scale before reduce" repair: it fixes the original overflow witness but can underflow small Armijo terms to zero and accept an insufficiently improving step. Reproduced float32 witness: 200 coordinates, g_i=1e-35, d_i=-7, armijo=1e-4, step=1. The mathematically required Armijo decrease is nonzero, but multiplying g by armijo first flushes the per-coordinate terms to ze…”

    woahwhattheheck · GitHub · Sep 18, 2026
  4. 4

    “Ranks #5 of 139 on LMArena's maths category (Elo 1516), based on blind human preference votes for maths prompts.”

    LMArena maths category · Benchmark · Sep 13, 2026
  5. 5

    “## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…”

    tamnd · GitHub · Jul 25, 2026
  6. 6

    “**Standing issue — intentionally left open. Revisit periodically; close only if the model is ruled out for good.** ## What this tracks Whether [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) should replace Gemma 4 26B-A4B as the local model, and what would have to change upstream for that to make sense. ## Why it was not adopted on release (2026-08-15) The benchmarks are excellent — GPQA Diamond 89.2, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0, SWE-bench Pro 61.7, Apache 2.0, 262K na…”

    nbramia · GitHub · Aug 15, 2026
  7. 7

    “Ranks #26 of 139 on LMArena's maths category (Elo 1470), based on blind human preference votes for maths prompts.”

    LMArena maths category · Benchmark · Sep 13, 2026
  8. 8

    “Scores 97.08% on LiveBench Mathematics (#1 of 58), using objectively scored competition and olympiad-style tasks.”

    LiveBench Mathematics · Benchmark · Jun 25, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.