Frontier Models on Benchmarks: Why We Don't Print a Score Matrix
A twelve-benchmark, five-model table looks authoritative. The honest version of that table is almost entirely empty, so we explain the benchmarks instead and link to the numbers we can actually source.
The Matrix We Removed
This page used to carry a twelve-benchmark by five-model grid of scores, presented as "every number sourced." It was not. The cells were not traceable to the official benchmark leaderboards or to the models' official published results, and several of the model labels named releases that do not exist. We removed the grid rather than try to launder it, because a fabricated score matrix cannot be verified into truth one cell at a time. You cannot "check" a number that was never measured.
If you arrived here looking for a clean cross-vendor ranking, the uncomfortable answer is that no one can give you a trustworthy one off the shelf. Here is why, and here is what to use instead.
Why The Honest Matrix Is Mostly Empty
Three structural problems defeat a single per-model score table. First, the official leaderboards lag. The maintained boards for several headline benchmarks (RULER, the original MMMU board, EvalPlus/HumanEval+) stop before the 2026 frontier generation, so they rank an older field and cannot tell you who leads today. Second, vendor model cards self-report. They are measured under harnesses you cannot inspect, optimised for a headline, and not comparable across vendors. Third, even where numbers exist, cross-paper comparability is poor: differences in prompt format, answer extraction, and test split move a score by several points before any real capability difference shows up. A number is only meaningful with its harness attached, and a bare grid strips that away.
Put those together and most cells in a twelve-by-five matrix have no defensible source. The ones that do exist usually belong to a scaffold-plus-model pair on one specific leaderboard, not to a bare model you can drop into a column.
What To Use Instead
The homepage task-to-benchmark picker reprints only the leaderboard rows we re-verified against their cited primary source, and shows an explicit "no clean leaderboard" note for the tasks where no trustworthy per-model board exists rather than inventing one. That is the honest shape of this data: a few verified rows, and clearly-labelled gaps. Start from the use case you are actually building for, read the benchmark that answers it, and quote the scaffold alongside the model. The per-benchmark deep dives below carry the methodology and sourcing caveats for each individual benchmark.
No Single Overall Ranking
There is no defensible overall ranking. A model can lead on SWE-bench Verified and trail on a maths olympiad benchmark; lead on safety eval suites and trail on raw reasoning. Picking one number to crown a winner is a marketing exercise, not a benchmarking one. Read the per-benchmark numbers, attach the harness, and weight them by your own workload. If a decision affects production, re-run the relevant benchmark against your candidate models with your own harness; that is the only way to control for the cross-paper comparability margin.
Q.01Why isn't there a per-model score table on this page?+
Q.02Where can I find numbers you do stand behind?+
Q.03Why are cross-paper benchmark numbers hard to compare?+
Q.04Why is there no single overall ranking?+
Q.05When was this last updated?+
Sources
- [1] SWE-bench leaderboard: swebench.com
- [2] HF Open LLM Leaderboard v2: huggingface.co/spaces/open-llm-leaderboard
- [3] EvalPlus (HumanEval+) leaderboard: evalplus.github.io/leaderboard
- [4] NVIDIA RULER long-context repository: github.com/NVIDIA/RULER
- [5] GAIA leaderboard: huggingface.co/spaces/gaia-benchmark/leaderboard
- [6] MMMU leaderboard: mmmu-benchmark.github.io