Recommendation for Competition math
Competition Math
Our top recommendation for Competition Math, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use when you need a model with validated competition math performance on both human preference and objective benchmarks, as it scores 95.99% on LiveBench Mathematics. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Deploy for competition math where LiveBench scores matter, as its 96.2% places it in the top tier of objectively evaluated models.
About this recommendation
- Updated
- Sep 25, 2026
- Evidence through
- Sep 25, 2026
- Sources
- 8
- Revision
- v78
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
3 of 6
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LiveBench Mathematics | 35% | #6 | 21/21 |
| LiveBench Reasoning | 25% | #10 | 21/21 |
| LMArena Math | 25% | #3 | 19/21 |
| LMArena Instruction Following | 10% | #4 | 20/21 |
| OpenRouter usage | 5% | 79/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic3 models
- OpenAI2 models
- Qwen1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 85 | 100% | no linked practitioner threads | #3 LMArena Math · #4 LMArena Instruction Following |
| 02 | GPT-5.6 SolOpenAI | 82 | 100% | 1 threads · 1 families · 1 cautions | #4 LiveBench Reasoning · #5 LiveBench Mathematics |
| 03 | Claude Opus 4.6Anthropic | 79 | 100% | no linked practitioner threads | #1 LMArena Instruction Following · #5 LMArena Math |
| 04 | GPT-5.6 LunaOpenAI | 78 | 100% | 1 threads · 1 families · 0 cautions | #25 LiveBench Reasoning · #34 LiveBench Mathematics |
| 05 | Qwen3.8 27BQwen | 75 | 100% | 1 threads · 1 families · 0 cautions | #25 LMArena Math · #37 LMArena Instruction Following |
| 06 | Claude Opus 5.5Anthropic | 63 | 44% | no linked practitioner threads | #1 LiveBench Mathematics · #2 LiveBench Reasoning |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Fable 5 places 4th in human preference rankings for math and 6th on objectively scored competition tasks, showing strong but not top-tier performance on rigorous olympiad problems.
Best when: Use when you need a model with validated competition math performance on both human preference and objective benchmarks, as it scores 95.99% on LiveBench Mathematics.
Tips
- Use when you need a model with validated competition math performance on both human preference and objective benchmarks, as it scores 95.99% on LiveBench Mathematics.
GPT-5.6 Sol achieves 96.2% on LiveBench Mathematics and ranks 5th, though an independent numerical review identified a floating-point underflow bug in its optimization routines that could affect certain mathematical derivations.
Best when: Deploy for competition math where LiveBench scores matter, as its 96.2% places it in the top tier of objectively evaluated models.
Tips
- Deploy for competition math where LiveBench scores matter, as its 96.2% places it in the top tier of objectively evaluated models.
Watch out for
- Watch for numerical stability issues in optimization-heavy derivations, as float32 underflow has been reproduced in Armijo condition calculations with small gradient terms.
Claude Opus 4.6 ranks 5th in human preference for math prompts but lacks objective benchmark scores, making its performance on rigorously scored competition problems harder to verify.
Best when: Consider when human-rated math helpfulness is your priority, as its 1516 Elo in LMArena's maths category reflects strong perceived quality in blind evaluations.
Tips
- Consider when human-rated math helpfulness is your priority, as its 1516 Elo in LMArena's maths category reflects strong perceived quality in blind evaluations.
GPT-5.6 Luna has demonstrated a complete, verified solve of a competition problem through an end-to-end proxy test, though no benchmark standings are available to compare against other models.
Best when: Use when you need a model with proven end-to-end solving capability on actual competition problems, as it passed a full slow-mode solve with dual-judge verification in under 5 minutes.
Tips
- Use when you need a model with proven end-to-end solving capability on actual competition problems, as it passed a full slow-mode solve with dual-judge verification in under 5 minutes.
Watch out for
- Note the lack of benchmark comparisons, as no LiveBench or LMArena scores are provided to gauge relative performance against competitors.
Qwen3.8 27B is an open-weight model with strong reported benchmarks including GPQA Diamond 89.2, though it ranks 26th in human preference for math and was not adopted upstream due to unresolved integration questions.
Best when: Choose for local deployment where Apache 2.0 licensing matters, as this 27B parameter model offers competitive benchmark scores without API dependencies.
Tips
- Choose for local deployment where Apache 2.0 licensing matters, as this 27B parameter model offers competitive benchmark scores without API dependencies.
Watch out for
- Expect lower human-rated math quality than top closed-weight alternatives, as its 1470 Elo in LMArena's maths category places it well outside the top tier.
Claude Opus 5.5 leads all models on LiveBench Mathematics with 97.08%, achieving the top rank on objectively scored competition and olympiad-style tasks.
Best when: Prioritize for maximum performance on rigorously evaluated competition math, as its 97.08% and #1 ranking on LiveBench Mathematics exceed all other reported scores.
Tips
- Prioritize for maximum performance on rigorously evaluated competition math, as its 97.08% and #1 ranking on LiveBench Mathematics exceed all other reported scores.
Frequently asked
- What is the top-ranked model for Competition Math?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use when you need a model with validated competition math performance on both human preference and objective benchmarks, as it scores 95.99% on LiveBench Mathematics.[1]
- What is an alternative to Anthropic: Claude Fable 5?
- OpenAI: GPT-5.6 Sol is the next-ranked option. Deploy for competition math where LiveBench scores matter, as its 96.2% places it in the top tier of objectively evaluated models.[2]
Sources
- 1
“Scores 95.99% on LiveBench Mathematics (#6 of 58), using objectively scored competition and olympiad-style tasks.”
LiveBench Mathematics · Benchmark · Jun 25, 2026 - 2
“Scores 96.2% on LiveBench Mathematics (#5 of 58), using objectively scored competition and olympiad-style tasks.”
LiveBench Mathematics · Benchmark · Jun 25, 2026 - 3
“INDEPENDENT NUMERICAL REVIEW · Z-Switch-Q9V3 / GPT-5.6 Sol · issue #15960 Fresh review found a blocker in the earlier proposed "scale before reduce" repair: it fixes the original overflow witness but can underflow small Armijo terms to zero and accept an insufficiently improving step. Reproduced float32 witness: 200 coordinates, g_i=1e-35, d_i=-7, armijo=1e-4, step=1. The mathematically required Armijo decrease is nonzero, but multiplying g by armijo first flushes the per-coordinate terms to ze…”
woahwhattheheck · GitHub · Sep 18, 2026 - 4
“Ranks #5 of 139 on LMArena's maths category (Elo 1516), based on blind human preference votes for maths prompts.”
LMArena maths category · Benchmark · Sep 13, 2026 - 5
“## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…”
tamnd · GitHub · Jul 25, 2026 - 6
“**Standing issue — intentionally left open. Revisit periodically; close only if the model is ruled out for good.** ## What this tracks Whether [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) should replace Gemma 4 26B-A4B as the local model, and what would have to change upstream for that to make sense. ## Why it was not adopted on release (2026-08-15) The benchmarks are excellent — GPQA Diamond 89.2, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0, SWE-bench Pro 61.7, Apache 2.0, 262K na…”
nbramia · GitHub · Aug 15, 2026 - 7
“Ranks #26 of 139 on LMArena's maths category (Elo 1470), based on blind human preference votes for maths prompts.”
LMArena maths category · Benchmark · Sep 13, 2026 - 8
“Scores 97.08% on LiveBench Mathematics (#1 of 58), using objectively scored competition and olympiad-style tasks.”
LiveBench Mathematics · Benchmark · Jun 25, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.