Our methodology

Trust the process.
Then inspect it.

This field guide is automated and non-commercial. That makes the method, limits, and audit trail more important, not less.

end-to-end process

Six gates from raw signal to a published answer

  1. 01

    Collect

    Catalog facts, benchmarks, prices, and public practitioner reports.

  2. 02

    Classify

    Associate evidence with the exact model, task, sentiment, and source type.

  3. 03

    Score

    Apply the use case's fixed weights to current, eligible model data.

  4. 04

    Explain

    Turn the fixed order and supplied evidence into readable recommendations.

  5. 05

    Verify

    Reject missing, mismatched, deleted, or unsupported citations.

  6. 06

    Publish

    Replace the complete recommendation atomically and rebuild the site.

Ranking, writing, and verification are separate stages. The writer never receives authority to reorder candidates, and a recommendation is not published unless every cited claim survives validation.

  1. 01

    Catalog, not copy

    OpenRouter and models.dev supply structured facts such as pricing, context, modalities, release dates, and capabilities. We never republish vendor descriptions.

  2. 02

    Evidence has a type

    Benchmarks, independent evaluators, practitioner reports, community discussions, and vendor-community posts are kept distinct. They carry different certainty and independence.

  3. 03

    The question sets the weights

    There is no universal best model. Each use case defines its own benchmark, price, context, popularity, and ownership signals. Only the newest coherent benchmark snapshots participate.

  4. 04

    Coverage affects confidence

    A model with one excellent result should not automatically equal one tested broadly. Scores receive a gentle confidence adjustment based on signal coverage.

  5. 05

    The writer does not rank

    The deterministic scorer owns candidate order. A language model turns supplied evidence into readable verdicts but cannot promote a model over the computed order.

  6. 06

    Every source is checked

    Citations are checked for model ownership and source existence, then a separate validation pass tests whether each excerpt supports the prose. One discussion thread counts as one independent source. Unsupported output fails closed.

  7. 07

    Freshness is explicit

    Catalog and benchmark jobs run several times daily; evidence gathering and synthesis run on their own schedules. Relevant evidence, material catalog changes, new benchmark snapshots, or the freshness ceiling trigger regeneration for affected use cases. A successful data job invokes the production deploy hook so the static site rebuilds from the new data.

known limits

A useful map is still not the territory.

Community evidence leans developer-heavy, public discussion can be promotional or mistaken, benchmark coverage varies, and listed API prices do not capture hosting or latency. Source families are capped and vendor-affiliated material is not treated as independent confirmation. A new model may appear in the catalog before independent evidence is strong enough to rank it. Treat the top pick as the strongest model to test first, then validate it against your prompts, latency target, and failure budget.