We do not print a per-model SOTA box for these benchmarks. The official boards (the ARC Prize leaderboard, the GPQA and HLE project pages) are JS-rendered and lag the 2026 frontier, and the vendor cards that do quote newer models self-report under harnesses we cannot inspect. What is durable is the structure: GPQA-Diamond, ARC-AGI-2, and Humanity's Last Exam all retain real headroom against their stated human baselines. For the numbers we are willing to stand behind, see the task picker on the homepage; for live scores, the official leaderboards linked in Sources below.
GPQA, ARC-AGI, and Humanity's Last Exam - The Reasoning Frontier 2026
MMLU is saturated. HumanEval is saturated. BIG-Bench Hard is approaching saturation. The benchmarks that still have real headroom in 2026 are GPQA-Diamond, ARC-AGI-2, and Humanity's Last Exam. These are the benchmarks frontier model comparisons should be made on.
Quick answer: what does a 75% GPQA-Diamond score mean?
GPQA-Diamond is the 198-question hardest subset of GPQA - PhD-written, multiple-choice questions in biology, chemistry, and physics that are deliberately Google-proof. A 75% score means a model answers roughly three-quarters of them correctly.
For reference, PhD experts in the relevant field score about 65% on the Diamond subset, and skilled non-experts with web access only about 34% (Rein et al., 2023). So 75% is above the expert baseline - the model is outperforming typical domain PhDs on this specific test - but it is no longer a frontier-leading result, because GPQA-Diamond scores have climbed steadily as the benchmark approaches saturation. Treat any single percentage as harness-dependent: the same model can move several points between zero-shot, chain-of-thought, and self-consistency settings.
GPQA - Graduate-Level Google-Proof Q&A
GPQA (Graduate-Level Google-Proof Q&A) was created by David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman (various institutions including New York University), published in 2023.
The benchmark contains 448 multiple-choice questions across biology, physics, and chemistry. Crucially, questions were written by PhD researchers with the explicit requirement that they be impossible to answer correctly by searching the web - you cannot Google your way to the right answer. Expert respondents (PhDs in the relevant field) answered correctly approximately 65% of the time on the Diamond subset. Non-experts with internet access answered correctly only 34% of the time.
GPQA-Diamond is the 198-question hardest subset, curated from the full 448 by keeping only questions where both expert annotators answered correctly (high expert agreement) yet the majority of skilled non-experts with web access still got them wrong. This is the canonical version for frontier model comparison. The expert PhD baseline of roughly 65% on the Diamond subset, reported in Rein et al. 2023, is the durable reference point: it is the line a model has to clear to claim above-expert performance.
We do not publish a per-model GPQA-Diamond score table here. A responsible cross-vendor table needs each cell sourced either to the official GPQA leaderboard or to a model's own published results under a disclosed harness, and for the 2026 frontier most of those cells have no defensible source: the maintained boards lag the current generation and vendor cards self-report under harnesses you cannot inspect. For the rows we are willing to stand behind, see the homepage task picker; for live numbers, the official GPQA materials linked in Sources.
ARC-AGI and ARC-AGI-2
ARC (Abstraction and Reasoning Corpus) was created by Francois Chollet at Google (2019) as a challenge to conventional AI benchmarking. Chollet's argument: intelligence should be measured by the ability to learn new skills from minimal examples, not by performance on tasks that can be memorised from training data. ARC tasks present a small number of input-output grid examples and ask the system to apply the inferred rule to a new input.
The 85% human baseline is the key reference point. Humans solve 85% of ARC tasks without prior exposure. Early frontier AI models scored near 0%. The benchmark became one of the most discussed AI challenges because it resisted every technique that worked on other benchmarks - CoT, tool use, fine-tuning.
In December 2024, OpenAI's o3 model scored 87.5% on the private ARC-AGI test set (above the human baseline), using extended compute (roughly $6,000 of inference per task at high-efficiency compute prices). This breakthrough triggered the development of ARC-AGI-2.
ARC-AGI-2 (Chollet, 2025) introduces significantly harder tasks designed to resist the high-compute, chain-of-thought techniques that cracked the original. It sits far below the human baseline, with real headroom. We do not quote a single per-model SOTA figure for it: the ARC Prize leaderboard is the live source of record, and verified-system scores there move faster than a static page can responsibly track.
Humanity's Last Exam
Humanity's Last Exam (HLE) was created by Scale AI and the Center for AI Safety (2025). The released public set contains roughly 2,500 questions written by domain experts across mathematics, sciences, and the humanities, specifically designed to remain hard for frontier models for several years. (The original paper described 3,000 questions; the set was trimmed after flagged and web-searchable items were removed.)
The name reflects a hypothesis: this will be the last benchmark on which human experts consistently outperform AI models. Frontier models still score well below the expert baseline, so the gap is substantial and the benchmark has enormous headroom. We do not print a single SOTA figure here: scores move with every release and the maintained boards lag the current generation. For the live numbers see the official HLE materials linked in Sources, and our HLE deep dive.
HLE is the benchmark to watch for frontier model comparisons over the next 2-3 years. When GPQA-Diamond saturates (expected within 12-18 months at current progress rates), HLE is the natural successor. Current frontier models still fail the large majority of expert-written questions - a useful reminder that even impressive models have significant gaps.
Why These Matter in 2026
The benchmark saturation problem is real and accelerating. A benchmark that can no longer discriminate between frontier models is useless for the most important use case - choosing a model for a specific deployment. In 2026, the benchmarks with real headroom are: GPQA-Diamond (7-8 point gap between frontier models), ARC-AGI-2 (15+ point gap), and Humanity's Last Exam (50+ point gap between frontier models and expert humans).
Any model comparison that does not include at least one of these three benchmarks is likely relying on saturated metrics. When a vendor card leads with MMLU, HumanEval, or BIG-Bench Hard, ask why they did not include GPQA-Diamond or ARC-AGI-2.
Frequently Asked Questions
What does ARC-AGI actually measure?+
Why is HLE called Humanity's Last Exam?+
What does a GPQA-Diamond score mean?+
What does a 75% GPQA-Diamond score mean?+
Is ARC-AGI a test of AGI?+
Sources
- [1] Rein et al., GPQA - arxiv.org/abs/2311.12022 - 2023
- [2] Chollet, ARC Challenge - arxiv.org/abs/1911.01547 - 2019
- [3] ARC Prize / ARC-AGI-2 leaderboard - arcprize.org
- [4] Humanity's Last Exam - agi.safe.ai/hle