Independent reference. No vendor affiliation. Reprinted scores dated to their source; elsewhere we link the official board. Editorial policy.
The Reasoning Frontier

We do not print a per-model SOTA box for these benchmarks. The official boards (the ARC Prize leaderboard, the GPQA and HLE project pages) are JS-rendered and lag the 2026 frontier, and the vendor cards that do quote newer models self-report under harnesses we cannot inspect. What is durable is the structure: GPQA-Diamond, ARC-AGI-2, and Humanity's Last Exam all retain real headroom against their stated human baselines. For the numbers we are willing to stand behind, see the task picker on the homepage; for live scores, the official leaderboards linked in Sources below.

Reviewed 2026

GPQA, ARC-AGI, and Humanity's Last Exam - The Reasoning Frontier 2026

MMLU is saturated. HumanEval is saturated. BIG-Bench Hard is approaching saturation. The benchmarks that still have real headroom in 2026 are GPQA-Diamond, ARC-AGI-2, and Humanity's Last Exam. These are the benchmarks frontier model comparisons should be made on.

Quick answer: what does a 75% GPQA-Diamond score mean?

GPQA-Diamond is the 198-question hardest subset of GPQA - PhD-written, multiple-choice questions in biology, chemistry, and physics that are deliberately Google-proof. A 75% score means a model answers roughly three-quarters of them correctly.

For reference, PhD experts in the relevant field score about 65% on the Diamond subset, and skilled non-experts with web access only about 34% (Rein et al., 2023). So 75% is above the expert baseline - the model is outperforming typical domain PhDs on this specific test - but it is no longer a frontier-leading result, because GPQA-Diamond scores have climbed steadily as the benchmark approaches saturation. Treat any single percentage as harness-dependent: the same model can move several points between zero-shot, chain-of-thought, and self-consistency settings.

GPQA - Graduate-Level Google-Proof Q&A

GPQA (Graduate-Level Google-Proof Q&A) was created by David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman (various institutions including New York University), published in 2023.

The benchmark contains 448 multiple-choice questions across biology, physics, and chemistry. Crucially, questions were written by PhD researchers with the explicit requirement that they be impossible to answer correctly by searching the web - you cannot Google your way to the right answer. Expert respondents (PhDs in the relevant field) answered correctly approximately 65% of the time on the Diamond subset. Non-experts with internet access answered correctly only 34% of the time.

GPQA-Diamond is the 198-question hardest subset, curated from the full 448 by keeping only questions where both expert annotators answered correctly (high expert agreement) yet the majority of skilled non-experts with web access still got them wrong. This is the canonical version for frontier model comparison. The expert PhD baseline of roughly 65% on the Diamond subset, reported in Rein et al. 2023, is the durable reference point: it is the line a model has to clear to claim above-expert performance.

We do not publish a per-model GPQA-Diamond score table here. A responsible cross-vendor table needs each cell sourced either to the official GPQA leaderboard or to a model's own published results under a disclosed harness, and for the 2026 frontier most of those cells have no defensible source: the maintained boards lag the current generation and vendor cards self-report under harnesses you cannot inspect. For the rows we are willing to stand behind, see the homepage task picker; for live numbers, the official GPQA materials linked in Sources.

ARC-AGI and ARC-AGI-2

ARC (Abstraction and Reasoning Corpus) was created by Francois Chollet at Google (2019) as a challenge to conventional AI benchmarking. Chollet's argument: intelligence should be measured by the ability to learn new skills from minimal examples, not by performance on tasks that can be memorised from training data. ARC tasks present a small number of input-output grid examples and ask the system to apply the inferred rule to a new input.

The 85% human baseline is the key reference point. Humans solve 85% of ARC tasks without prior exposure. Early frontier AI models scored near 0%. The benchmark became one of the most discussed AI challenges because it resisted every technique that worked on other benchmarks - CoT, tool use, fine-tuning.

In December 2024, OpenAI's o3 model scored 87.5% on the private ARC-AGI test set (above the human baseline), using extended compute (roughly $6,000 of inference per task at high-efficiency compute prices). This breakthrough triggered the development of ARC-AGI-2.

ARC-AGI-2 (Chollet, 2025) introduces significantly harder tasks designed to resist the high-compute, chain-of-thought techniques that cracked the original. It sits far below the human baseline, with real headroom. We do not quote a single per-model SOTA figure for it: the ARC Prize leaderboard is the live source of record, and verified-system scores there move faster than a static page can responsibly track.

Humanity's Last Exam

Humanity's Last Exam (HLE) was created by Scale AI and the Center for AI Safety (2025). The released public set contains roughly 2,500 questions written by domain experts across mathematics, sciences, and the humanities, specifically designed to remain hard for frontier models for several years. (The original paper described 3,000 questions; the set was trimmed after flagged and web-searchable items were removed.)

The name reflects a hypothesis: this will be the last benchmark on which human experts consistently outperform AI models. Frontier models still score well below the expert baseline, so the gap is substantial and the benchmark has enormous headroom. We do not print a single SOTA figure here: scores move with every release and the maintained boards lag the current generation. For the live numbers see the official HLE materials linked in Sources, and our HLE deep dive.

HLE is the benchmark to watch for frontier model comparisons over the next 2-3 years. When GPQA-Diamond saturates (expected within 12-18 months at current progress rates), HLE is the natural successor. Current frontier models still fail the large majority of expert-written questions - a useful reminder that even impressive models have significant gaps.

Why These Matter in 2026

The benchmark saturation problem is real and accelerating. A benchmark that can no longer discriminate between frontier models is useless for the most important use case - choosing a model for a specific deployment. In 2026, the benchmarks with real headroom are: GPQA-Diamond (7-8 point gap between frontier models), ARC-AGI-2 (15+ point gap), and Humanity's Last Exam (50+ point gap between frontier models and expert humans).

Any model comparison that does not include at least one of these three benchmarks is likely relying on saturated metrics. When a vendor card leads with MMLU, HumanEval, or BIG-Bench Hard, ask why they did not include GPQA-Diamond or ARC-AGI-2.

Frequently Asked Questions

What does ARC-AGI actually measure?+
ARC-AGI measures the ability to induce abstract rules from visual examples and apply them to new cases. Each task shows a few input-output grid pairs, and the model must identify the underlying transformation rule and apply it to a new input. Humans find these tasks straightforward (average 85%). ARC-AGI-2 is a harder version designed to resist the techniques that cracked the original in December 2024.
Why is HLE called Humanity's Last Exam?+
Humanity's Last Exam is named for the hypothesis that it will be the last benchmark where human experts consistently outperform frontier AI. The roughly 2,500 questions in the released set were written by domain experts specifically to be hard for AI (the original paper described 3,000, trimmed after flagged and web-searchable items were removed). Current models score well below the expert baseline, with several years of expected headroom.
What does a GPQA-Diamond score mean?+
GPQA-Diamond tests graduate-level scientific reasoning across biology, chemistry, and physics. A 70% GPQA-Diamond score means the model correctly answers 70% of questions that stump most domain experts. Human expert baseline is approximately 65% on GPQA-Diamond, meaning frontier models now exceed typical PhD-level performance on this specific benchmark.
What does a 75% GPQA-Diamond score mean?+
A 75% GPQA-Diamond score means a model answers about 75% of the 198 PhD-written, Google-proof science questions correctly. The expert baseline is roughly 65% (PhDs in the relevant field) and non-experts with web access score about 34% (Rein et al., 2023), so 75% is clearly above expert level - though no longer frontier-leading, as GPQA-Diamond scores have climbed as the benchmark approaches saturation. Treat any single percentage as harness-dependent.
Is ARC-AGI a test of AGI?+
ARC-AGI measures a specific capability (visual abstract reasoning with minimal training examples) that its creator argues is more aligned with general intelligence than knowledge recall. Scoring 85% on ARC-AGI does not mean a system is generally intelligent. The benchmark is valuable for measuring this narrow but challenging capability.

Sources

  1. [1] Rein et al., GPQA - arxiv.org/abs/2311.12022 - 2023
  2. [2] Chollet, ARC Challenge - arxiv.org/abs/1911.01547 - 2019
  3. [3] ARC Prize / ARC-AGI-2 leaderboard - arcprize.org
  4. [4] Humanity's Last Exam - agi.safe.ai/hle
From the editor

Benchmarking Agents Review is published by Digital Signet, an independent firm that builds and ships AI agents in production for mid-market companies. If you are evaluating, designing, or productionising LLM agents and want a working second opinion, get in touch.

Book a 30-min scoping callDigital Signet →

30 minutes, free, independent.·1-page action plan within 48h.·Honest if not the right fit.