Recommendation for Grounded QA

Grounded RAG

The best LLMs for grounded RAG cannot be determined from the available evidence, which only provides general preference rankings.[1][2][3][3] The evidence shows Claude Opus 4.6 ranks #1 (Elo 1512) on LMArena's text arena, and Claude Fable 5 ranks #2 (Elo 1504), but these are blind human preference votes, not groundedness benchmarks. The only RAG-related evidence describes a model selection methodology for a bombing-related question, listing GPT-5.4 and Claude Opus 4.7 as parametric models tested against retrieval-augmented ones like Gemini 3 Pro + Search and Sonar Pro. It also notes that two search-equipped models disagreed with each other. None of this measures hallucination rates, citation accuracy, or honest refusals. For grounded RAG, you will need to evaluate models on domain-specific retrieval tasks yourself.

About this recommendation

Updated
Jul 21, 2026
Evidence through
Jul 21, 2026
Sources
8
Revision
v2
  1. Muse Spark 1.1 ranks #7 (Elo 1491), landing it in the middle of the pack. Without RAG-specific benchmarks, its groundedness remains unknown.

    Best when: When you are evaluating multiple mid-tier models and need to benchmark grounding performance yourself.

    Tips

    • Ranks #7 of 17, placing it in the upper half of the arena leaderboard.
      Source 4
      Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • No available evidence on retrieval augmentation, source faithfulness, or refusal behavior.
      Source 4
      Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  2. Claude Opus 4.6 holds the top spot on LMArena's text arena with an Elo of 1512. However, this ranking reflects general human preference, not grounded RAG performance. You should still verify citation accuracy and hallucination rates for your specific domain.

    Best when: When you want a model that human evaluators consistently prefer for general text quality, though grounding performance remains unverified.

    Tips

    • Ranks #1 of 17 on LMArena's overall text arena, indicating strong general text quality as judged by blind human preference votes.
      Source 1
      Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • No evidence available about groundedness, citation accuracy, or hallucination rates, so you must test these yourself before deployment.
      Source 1
      Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  3. Claude Fable 5 ranks #2 with an Elo of 1504 on LMArena's text arena. Like other high performers, its grounded RAG capabilities are not demonstrated in the available evidence.

    Best when: When you need a high-ranking model by human preference standards, but you have your own evaluation pipeline for RAG-specific tasks.

    Tips

    • Ranks #2 of 17 on LMArena's overall text arena with strong preference scores from blind human evaluators.
      Source 2
      Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Ranking based on general preference, not RAG-specific metrics like citation fidelity or refusal accuracy.
      Source 2
      Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  4. Claude Opus 4.7 ranks #4 (Elo 1499) and was tested in a research setting as a parametric model. Like GPT-5.4, it was compared against retrieval-augmented alternatives.

    Best when: When you want a parametric model with high arena ranking for custom RAG pipeline development.

    Tips

    • Selected as a frontier model for research comparing parametric and retrieval-augmented capability surfaces.
      Source 3
      Fwiw the two models that did have access to search disagreed with each other on the bombing one: > 7.1 Model selection > Five frontier models, chosen to cover two capability surfaces: > Parametric (training-only): GPT-5.4 (OpenAI), Claude Opus 4.7 (Anthropic), Gemini 3 Pro (Google) > Retrieval-augmented: Gemini 3 Pro + Search (Google), Sonar Pro (Perplexity)
    • Ranks #4 of 17 on LMArena's overall text arena with an Elo of 1499.
      Source 5
      Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Classified as parametric (training-only), so you must supply your own retrieval infrastructure.
      Source 3
      Fwiw the two models that did have access to search disagreed with each other on the bombing one: > 7.1 Model selection > Five frontier models, chosen to cover two capability surfaces: > Parametric (training-only): GPT-5.4 (OpenAI), Claude Opus 4.7 (Anthropic), Gemini 3 Pro (Google) > Retrieval-augmented: Gemini 3 Pro + Search (Google), Sonar Pro (Perplexity)
  5. GPT-5.4 ranks #3 (Elo 1499) and appears in a research comparison as a parametric (training-only) model alongside Claude Opus 4.7. The study contrasted it against retrieval-augmented models like Gemini 3 Pro + Search and Sonar Pro.

    Best when: When you need a parametric baseline model for comparison against retrieval-augmented systems in your own benchmarks.

    Tips

    • Included in formal research methodology as a frontier parametric model alongside Claude Opus 4.7 and Gemini 3 Pro.
      Source 3
      Fwiw the two models that did have access to search disagreed with each other on the bombing one: > 7.1 Model selection > Five frontier models, chosen to cover two capability surfaces: > Parametric (training-only): GPT-5.4 (OpenAI), Claude Opus 4.7 (Anthropic), Gemini 3 Pro (Google) > Retrieval-augmented: Gemini 3 Pro + Search (Google), Sonar Pro (Perplexity)
    • Ranks #3 of 17 on LMArena's overall text arena with an Elo of 1499.
      Source 6
      Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Categorized as parametric (training-only), meaning it lacks built-in retrieval and requires external systems for grounding.
      Source 3
      Fwiw the two models that did have access to search disagreed with each other on the bombing one: > 7.1 Model selection > Five frontier models, chosen to cover two capability surfaces: > Parametric (training-only): GPT-5.4 (OpenAI), Claude Opus 4.7 (Anthropic), Gemini 3 Pro (Google) > Retrieval-augmented: Gemini 3 Pro + Search (Google), Sonar Pro (Perplexity)
  6. GPT-5.4 Mini ranks #5 with an Elo of 1496. This smaller variant may offer cost or latency benefits, but the evidence provides no RAG-specific data.

    Best when: When you need a lighter-weight model with arena-validated general quality and plan to handle grounding through external retrieval.

    Tips

    • Ranks #5 of 17 on LMArena's overall text arena, placing it in the top tier by preference votes.
      Source 7
      Ranks #5 of 17 on LMArena's overall text arena (Elo 1496), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • No evidence about retrieval capabilities, citation behavior, or hallucination tendencies for grounded RAG workloads.
      Source 7
      Ranks #5 of 17 on LMArena's overall text arena (Elo 1496), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  7. GPT-5.5 ranks #6 with an Elo of 1491, just behind GPT-5.4 Mini. Arena ranking alone does not predict RAG faithfulness.

    Best when: When general text quality matters and you have your own pipeline to enforce source fidelity.

    Tips

    • Ranks #6 of 17 on LMArena's overall text arena, indicating competitive performance in blind preference evaluations.
      Source 8
      Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Arena preference votes measure general appeal, not groundedness or citation accuracy.
      Source 8
      Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗

Frequently asked

Which model ranks highest on LMArena?
Claude Opus 4.6 ranks #1 of 17 with an Elo of 1512, followed by Claude Fable 5 at #2 (Elo 1504), based on blind human preference votes.[1][2]
What models were used in the parametric vs retrieval-augmented comparison?
Parametric models included GPT-5.4, Claude Opus 4.7, and Gemini 3 Pro, while retrieval-augmented models were Gemini 3 Pro + Search and Sonar Pro.[3][3]
Do search-enabled models agree with each other?
Not necessarily. The evidence notes that two models with search access disagreed with each other on a bombing-related question.[3][3]

Sources

  1. 1

    Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  2. 2

    Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  3. 3

    Fwiw the two models that did have access to search disagreed with each other on the bombing one: > 7.1 Model selection > Five frontier models, chosen to cover two capability surfaces: > Parametric (training-only): GPT-5.4 (OpenAI), Claude Opus 4.7 (Anthropic), Gemini 3 Pro (Google) > Retrieval-augmented: Gemini 3 Pro + Search (Google), Sonar Pro (Perplexity)

    post-it · Hacker News · May 28, 2026
  4. 4

    Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  5. 5

    Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  6. 6

    Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  7. 7

    Ranks #5 of 17 on LMArena's overall text arena (Elo 1496), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  8. 8

    Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.