Recommendation for Grounded QA
Grounded RAG
The best LLMs for grounded RAG cannot be determined from the available evidence, which only provides general preference rankings.[1][2][3][3] The evidence shows Claude Opus 4.6 ranks #1 (Elo 1512) on LMArena's text arena, and Claude Fable 5 ranks #2 (Elo 1504), but these are blind human preference votes, not groundedness benchmarks. The only RAG-related evidence describes a model selection methodology for a bombing-related question, listing GPT-5.4 and Claude Opus 4.7 as parametric models tested against retrieval-augmented ones like Gemini 3 Pro + Search and Sonar Pro. It also notes that two search-equipped models disagreed with each other. None of this measures hallucination rates, citation accuracy, or honest refusals. For grounded RAG, you will need to evaluate models on domain-specific retrieval tasks yourself.
About this recommendation
- Updated
- Jul 21, 2026
- Evidence through
- Jul 21, 2026
- Sources
- 8
- Revision
- v2
Muse Spark 1.1 ranks #7 (Elo 1491), landing it in the middle of the pack. Without RAG-specific benchmarks, its groundedness remains unknown.
Best when: When you are evaluating multiple mid-tier models and need to benchmark grounding performance yourself.
Tips
- Ranks #7 of 17, placing it in the upper half of the arena leaderboard.
Watch out for
- No available evidence on retrieval augmentation, source faithfulness, or refusal behavior.
Claude Opus 4.6 holds the top spot on LMArena's text arena with an Elo of 1512. However, this ranking reflects general human preference, not grounded RAG performance. You should still verify citation accuracy and hallucination rates for your specific domain.
Best when: When you want a model that human evaluators consistently prefer for general text quality, though grounding performance remains unverified.
Tips
- Ranks #1 of 17 on LMArena's overall text arena, indicating strong general text quality as judged by blind human preference votes.
Watch out for
- No evidence available about groundedness, citation accuracy, or hallucination rates, so you must test these yourself before deployment.
Claude Fable 5 ranks #2 with an Elo of 1504 on LMArena's text arena. Like other high performers, its grounded RAG capabilities are not demonstrated in the available evidence.
Best when: When you need a high-ranking model by human preference standards, but you have your own evaluation pipeline for RAG-specific tasks.
Tips
- Ranks #2 of 17 on LMArena's overall text arena with strong preference scores from blind human evaluators.
Watch out for
- Ranking based on general preference, not RAG-specific metrics like citation fidelity or refusal accuracy.
Claude Opus 4.7 ranks #4 (Elo 1499) and was tested in a research setting as a parametric model. Like GPT-5.4, it was compared against retrieval-augmented alternatives.
Best when: When you want a parametric model with high arena ranking for custom RAG pipeline development.
Tips
- Selected as a frontier model for research comparing parametric and retrieval-augmented capability surfaces.
- Ranks #4 of 17 on LMArena's overall text arena with an Elo of 1499.
Watch out for
- Classified as parametric (training-only), so you must supply your own retrieval infrastructure.
GPT-5.4 ranks #3 (Elo 1499) and appears in a research comparison as a parametric (training-only) model alongside Claude Opus 4.7. The study contrasted it against retrieval-augmented models like Gemini 3 Pro + Search and Sonar Pro.
Best when: When you need a parametric baseline model for comparison against retrieval-augmented systems in your own benchmarks.
Tips
- Included in formal research methodology as a frontier parametric model alongside Claude Opus 4.7 and Gemini 3 Pro.
- Ranks #3 of 17 on LMArena's overall text arena with an Elo of 1499.
Watch out for
- Categorized as parametric (training-only), meaning it lacks built-in retrieval and requires external systems for grounding.
GPT-5.4 Mini ranks #5 with an Elo of 1496. This smaller variant may offer cost or latency benefits, but the evidence provides no RAG-specific data.
Best when: When you need a lighter-weight model with arena-validated general quality and plan to handle grounding through external retrieval.
Tips
- Ranks #5 of 17 on LMArena's overall text arena, placing it in the top tier by preference votes.
Watch out for
- No evidence about retrieval capabilities, citation behavior, or hallucination tendencies for grounded RAG workloads.
GPT-5.5 ranks #6 with an Elo of 1491, just behind GPT-5.4 Mini. Arena ranking alone does not predict RAG faithfulness.
Best when: When general text quality matters and you have your own pipeline to enforce source fidelity.
Tips
- Ranks #6 of 17 on LMArena's overall text arena, indicating competitive performance in blind preference evaluations.
Watch out for
- Arena preference votes measure general appeal, not groundedness or citation accuracy.
Frequently asked
Sources
- 1
“Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 2
“Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 3
“Fwiw the two models that did have access to search disagreed with each other on the bombing one: > 7.1 Model selection > Five frontier models, chosen to cover two capability surfaces: > Parametric (training-only): GPT-5.4 (OpenAI), Claude Opus 4.7 (Anthropic), Gemini 3 Pro (Google) > Retrieval-augmented: Gemini 3 Pro + Search (Google), Sonar Pro (Perplexity)”
post-it · Hacker News · May 28, 2026 - 4
“Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 5
“Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 6
“Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 7
“Ranks #5 of 17 on LMArena's overall text arena (Elo 1496), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 8
“Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.