Independent reference. No vendor affiliation. Reprinted scores dated to their source; elsewhere we link the official board. Editorial policy.
Abstract
WhatWhy this page does not reprint a per-model benchmark score matrix, and where the numbers we trust live instead.
The problemCross-vendor score cells mostly can't be sourced to a primary leaderboard for the current frontier.
What we trustThe verified per-task leaderboards in the homepage picker, plus the per-benchmark deep dives.
ReviewedJune 2026.
Section III Frontier Reference|Reviewed 2026

Frontier Models on Benchmarks: Why We Don't Print a Score Matrix

A twelve-benchmark, five-model table looks authoritative. The honest version of that table is almost entirely empty, so we explain the benchmarks instead and link to the numbers we can actually source.

I

The Matrix We Removed

This page used to carry a twelve-benchmark by five-model grid of scores, presented as "every number sourced." It was not. The cells were not traceable to the official benchmark leaderboards or to the models' official published results, and several of the model labels named releases that do not exist. We removed the grid rather than try to launder it, because a fabricated score matrix cannot be verified into truth one cell at a time. You cannot "check" a number that was never measured.

If you arrived here looking for a clean cross-vendor ranking, the uncomfortable answer is that no one can give you a trustworthy one off the shelf. Here is why, and here is what to use instead.

II

Why The Honest Matrix Is Mostly Empty

Three structural problems defeat a single per-model score table. First, the official leaderboards lag. The maintained boards for several headline benchmarks (RULER, the original MMMU board, EvalPlus/HumanEval+) stop before the 2026 frontier generation, so they rank an older field and cannot tell you who leads today. Second, vendor model cards self-report. They are measured under harnesses you cannot inspect, optimised for a headline, and not comparable across vendors. Third, even where numbers exist, cross-paper comparability is poor: differences in prompt format, answer extraction, and test split move a score by several points before any real capability difference shows up. A number is only meaningful with its harness attached, and a bare grid strips that away.

Put those together and most cells in a twelve-by-five matrix have no defensible source. The ones that do exist usually belong to a scaffold-plus-model pair on one specific leaderboard, not to a bare model you can drop into a column.

III

What To Use Instead

The homepage task-to-benchmark picker reprints only the leaderboard rows we re-verified against their cited primary source, and shows an explicit "no clean leaderboard" note for the tasks where no trustworthy per-model board exists rather than inventing one. That is the honest shape of this data: a few verified rows, and clearly-labelled gaps. Start from the use case you are actually building for, read the benchmark that answers it, and quote the scaffold alongside the model. The per-benchmark deep dives below carry the methodology and sourcing caveats for each individual benchmark.

IV

No Single Overall Ranking

There is no defensible overall ranking. A model can lead on SWE-bench Verified and trail on a maths olympiad benchmark; lead on safety eval suites and trail on raw reasoning. Picking one number to crown a winner is a marketing exercise, not a benchmarking one. Read the per-benchmark numbers, attach the harness, and weight them by your own workload. If a decision affects production, re-run the relevant benchmark against your candidate models with your own harness; that is the only way to control for the cross-paper comparability margin.

Cost calculatorSWE-bench deep diveHumanity's Last Exam
Reader Questions
Q.01Why isn't there a per-model score table on this page?+
Because we could not source the cells to primary leaderboards we trust. A cross-vendor benchmark matrix needs each cell taken either from the benchmark's official leaderboard or the model's official published results, under a disclosed harness, for an actually-released model. When you try to fill 60 cells that way you find most of them have no defensible source: the official boards lag the current frontier, vendor cards self-report under harnesses you can't inspect, and cross-paper numbers carry a several-point comparability margin. We would rather publish nothing than publish a matrix we cannot stand behind.
Q.02Where can I find numbers you do stand behind?+
The homepage task-to-benchmark picker reprints only leaderboard rows we re-verified against the cited primary source (the official SWE-bench data file, EvalPlus results.json, the NVIDIA RULER repository table), and it shows an honest 'no clean leaderboard' note for tasks where no trustworthy per-model board exists. Each per-benchmark deep dive linked below walks through the methodology and the sourcing caveats for that specific benchmark.
Q.03Why are cross-paper benchmark numbers hard to compare?+
Different evaluations use different harnesses, prompt formats, answer-extraction rules, and sometimes different test splits. Two papers reporting 'GPQA Diamond' can differ by several points on formatting and scoring alone, before any difference in the underlying model. A score is only meaningful with its harness attached, which is why a bare matrix of numbers misleads more than it informs.
Q.04Why is there no single overall ranking?+
There is no defensible overall ranking. A model can lead on SWE-bench Verified and trail on a maths olympiad benchmark. A model can lead on safety eval suites and trail on raw reasoning. Picking one number to declare a winner is a marketing exercise, not a benchmarking one. Read the per-benchmark numbers and weight them by your actual use case.
Q.05When was this last updated?+
Reviewed June 2026. The verified leaderboard rows in the homepage picker carry their own capture date and are re-pulled from primary sources; this page links to them rather than copying numbers that would go stale.

Sources

  1. [1] SWE-bench leaderboard: swebench.com
  2. [2] HF Open LLM Leaderboard v2: huggingface.co/spaces/open-llm-leaderboard
  3. [3] EvalPlus (HumanEval+) leaderboard: evalplus.github.io/leaderboard
  4. [4] NVIDIA RULER long-context repository: github.com/NVIDIA/RULER
  5. [5] GAIA leaderboard: huggingface.co/spaces/gaia-benchmark/leaderboard
  6. [6] MMMU leaderboard: mmmu-benchmark.github.io
From the editor

Benchmarking Agents Review is published by Digital Signet, an independent firm that builds and ships AI agents in production for mid-market companies. If you are evaluating, designing, or productionising LLM agents and want a working second opinion, get in touch.

Book a 30-min scoping callDigital Signet →

30 minutes, free, independent.·1-page action plan within 48h.·Honest if not the right fit.