Recommendation for Best value
Best Value LLM
The best value LLM right now is GPT-5.5, offering near-frontier capability at a fraction of flagship pricing. According to community reports, it delivers roughly 90% of GPT-5.5's performance at about 20% of the cost for coding tasks, making it a compelling daily driver for developers watching their spend. GPT-5.6 Luna and Terra also deserve attention for value-conscious builders. Both outperform Claude Fable 5 on Agents' Last Exam at roughly one-sixth the estimated cost, according to OpenAI's published benchmarks. This efficiency tier fills an important gap between budget options and the ultra-premium tier occupied by Fable and full-fat reasoning models. GLM 5.2 is harder to recommend here, with community feedback citing instruction-following issues and a cost-per-task that rivals significantly stronger models like GPT-5.6 Sol Max. For pure dollar-efficiency on agentic benchmarks, smaller specialized models may win, but GPT-5.5 and the GPT-5.6 mid-tier strike the best balance of usable quality and sustainable pricing for production workloads.
About this recommendation
- Updated
- Jul 17, 2026
- Evidence through
- Jul 17, 2026
- Sources
- 8
- Revision
- v1
GPT-5.5 is the clear value leader, delivering near-frontier reasoning at a price point that makes it viable as a daily workhorse rather than a specialty tool. Users report it handles 90% of what the flagship models can do at 20% of the cost, which is precisely what capability-per-dollar should mean.
Best when: You need strong coding and reasoning performance daily without burning through a monthly budget cap.
Tips
- Achieves roughly 90% of flagship capability at about 20% of the cost for coding tasks.
- Four times better reasoning efficiency than Claude Opus while priced lower.
- Generates fewer tokens than Opus 4.8, lowering net cost despite higher per-token rates.
- Ranks #10 on LMArena overall text arena with strong human preference scores.
Watch out for
- Still expensive enough that some users cannot justify it as a main model under workplace spend caps.
- Token caps and throughput have fluctuated significantly according to subscription tracking.
GPT-5.6 Luna earns its spot by beating Claude Fable 5 on a rigorous agentic benchmark at roughly one-sixth the cost. It fills the sweet spot for builders who want serious capability without paying the reasoning-model premium.
Best when: You want strong agentic performance for long workflows and care about cost efficiency at scale.
Tips
- Outperforms Claude Fable 5 on Agents' Last Exam at approximately one-sixth the estimated cost.
- Part of a smaller-model efficiency tier designed to make intelligence more affordable.
Watch out for
- Lacks independent community validation of the efficiency claims beyond OpenAI's own benchmarks.
GPT-5.6 Terra matches Luna's value proposition with the same benchmark efficiency claim against Fable 5. It is a solid alternative if your workload benefits from this particular model variant's tuning.
Best when: You need a mid-tier model with proven agentic efficiency at the same cost level as Luna.
Tips
- Matches Luna in outperforming Claude Fable 5 at roughly one-sixth the cost.
- Designed specifically to make capable intelligence more abundant and affordable.
Watch out for
- Benchmark claims come solely from OpenAI with no third-party cost validation yet.
GLM 5.2 looks competitive on paper but struggles in practice. The cost-per-task lands awkwardly close to stronger models while instruction-following issues undermine its real-world value.
Best when: You need agentic capabilities and are willing to tolerate weaker instruction adherence for the price.
Tips
- Scores well on agentic benchmarks despite being from a non-major provider.
- GLM 5.2 Max variant costs roughly half the per-task price of GPT-5.6 Sol Max.
Watch out for
- Instruction following falls short of expectations for reliable daily use.
- Cost per task of $0.94 is nearly identical to GPT-5.6 Sol Max despite being weaker.
- Only 744b parameters makes it less competitive than newer offerings.
GPT-5.6 Sol is an efficiency benchmark standout, beating Claude Fable 5 by over 11 points at one-quarter the cost. It lands here rather than higher because Sol targets maximum capability rather than pure value, pricing itself above the mid-tier sweet spot.
Best when: You want top-tier agentic performance with better efficiency than the most expensive reasoning models.
Tips
- Scores 53.6 on Agents' Last Exam, beating Claude Fable 5 by 13.1 points.
- At medium reasoning, beats Fable 5 by 11.4 points at roughly one-quarter the cost.
Watch out for
- Targets flagship capability tier, not the budget-conscious value segment.
Muse Spark 1.1 ranks #4 on LMArena, which suggests strong quality, but the total absence of pricing or efficiency data makes it impossible to evaluate for value. Could be a contender if pricing is revealed to be competitive.
Best when: You want strong preference-ranked quality and cost is not a primary constraint.
Tips
- Ranks #4 of 49 on LMArena with an Elo of 1481 based on blind human votes.
Watch out for
- No pricing or efficiency data available to assess value proposition.
Frequently asked
- Which model gives the best quality per dollar?
- GPT-5.5 is widely cited as 90% as capable as flagship models at roughly 20% of the cost, making it the top value pick for coding and general tasks.
- Is GPT-5.5 cheaper than Claude Opus 4.8?
- Yes. GPT-5.5 generates fewer tokens than Opus 4.8, and despite higher per-token pricing, the net cost ends up lower because it is less verbose.
- Are the GPT-5.6 smaller models good value?
- GPT-5.6 Luna and Terra both outperform Claude Fable 5 on Agents' Last Exam at around one-sixth the cost, according to OpenAI's benchmark data.
- How does GLM 5.2 compare on price?
- GLM 5.2 costs about $0.94 per task, similar to GPT-5.6 Sol Max, but users report weaker instruction following, hurting its value proposition.
Sources
- 1
“I prefer GPT 5.5 to Opus but both are absurdly expensive token hogs, I can't afford to use either as my main model at $work with the monthly spend cap we have. I use Composer (since we use Cursor) or GPT 5.3-codex as my workhorse models and only break out the big guns when I have a genuinely difficult problem to solve. IMO somewhat weirdly 5.3-codex might be the best overall coding model OpenAI have ever released. It's 90% as good as 5.5 and costs about 20% as much, since it's both cheaper per…”
ifwinterco · Hacker News · Jun 30, 2026 - 2
“It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2 $6. For comparison, GPT 5.4 is $2.5 $15, GPT 5.5 5.6 are $5 $30, Opus 4.8 is $5 $25, Fable is $10 $50. And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https: x.com elonmusk status 2074911038286295049. I guess the Cursor data was very useful.”
Tiberium · Hacker News · Jul 8, 2026 - 3
“Most variants of GPT-5.5 are less chatty and token-intensive than Opus 4.8 4.7, so despite the output token price being higher, it generates fewer tokens, so the net cost is lower. Per-token pricing is totally sensible from the provider-perspective on mapping COGS to revenue, but for a consumer, different models will produce more or less tokens, meaning the cost calculation is multi-dimensional.”
Spartan-S63 · Hacker News · May 28, 2026 - 4
“Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 5
“Building a tool to compare value across LLM provider options. Part of it tracks how many tokens you actually get from various subscriptions, over time. Past week, multiple people asked me about it — they'd been hitting Claude and Codex limits faster than expected. Ran the tests yesterday. Reran today. Here's what came back: ▸ ChatGPT Plus GPT-5.5: 95M → 37M tokens week (−61%) ▸ Claude Max 20× Sonnet 4.6: 388M → 214M (−45%) ▸ Claude Max 20× Opus 4.7: 248M → 162M (−35%) ▸ Claude Pro Sonnet 4.6: 1…”
wonderwhyer · Hacker News · Apr 30, 2026 - 6
“"On Agents’ Last Exam (opens in a new window), an evaluation of long-running professional workflows across 55 fields, GPT‑5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. That efficiency extends to smaller models, which are essential to making intelligence more abundant and affordable: GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-s…”
saberience · Hacker News · Jul 9, 2026 - 7
“Wow, seems worse even on price performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberGym"”
conradkay · Hacker News · Jun 30, 2026 - 8
“Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.