Section I.iv Knowledge and Reasoning|Reviewed 2026
MATH Benchmark: 12,500 Problems, 99%+ Saturation, What Replaced It
The benchmark that drove the chain-of-thought era, now functionally retired at the top.
Which mathematics benchmark should I use in 2026?For a headline frontier-math comparison, quote AIME 2025 (hard and discriminating, though high-variance at 15 problems per year) alongside FrontierMath (the hardest math benchmark in active use, with real headroom on its Tier 4 expansion set). GSM8K and MATH are both saturated above the high-90s and now serve only as sanity checks or mid-tier and open-weight comparisons. The discriminating signal in 2026 lives in AIME and FrontierMath, not the older benchmarks.
The mathematics-benchmark landscape in 2026
"Mathematics benchmark" is not one test. The five that matter in 2026 span grade-school word problems through unsolved research mathematics, and most of the older ones are saturated. The table below maps each to what it measures and where it is still useful; the rest of this page is the deep dive on MATH itself, the benchmark that defined the category.
Benchmark
What it measures
2026 status
Use it for
GSM8K
8,500 grade-school word problems (Cobbe et al., 2021)
Saturated (frontier high-90s)
Legacy sanity check
MATH
12,500 competition problems (Hendrycks et al., 2021)
Saturated (>99% on MATH-500)
Sanity check, mid-tier
AIME 2024/2025
15 contest problems/year, integer answers 0-999
Discriminating, high variance
Headline frontier comparison
FrontierMath
338 research-grade problems (Epoch AI, v2 Jun 2026)
Hardest in active use; Tier 4 unsaturated
Frontier reasoning ceiling
Putnam-Bench
Putnam-level proof problems
Retains headroom
Olympiad / proof reasoning
I
The Construction
MATH was published in March 2021 by Dan Hendrycks and collaborators at UC Berkeley. It draws 12,500 problems from American math competitions (AMC 10, AMC 12, AIME, USAMO qualifiers, and similar). Problems are split 7,500 train / 5,000 test, tagged by subject and by difficulty level 1 through 5. Each ships with a full LaTeX solution.
The grading rule is exact match on the final boxed answer after the model emits its full chain-of-thought. There is no partial credit, no judge model, no human in the loop at scoring time. This makes MATH cheap to run and reproducible, two reasons it became the de facto math benchmark for three years.
II
SOTA Progression 2021 to 2026
Date
Tier / Score
Note
Mar 2021
GPT-3 baseline at 6.9%
Original MATH paper, no CoT, no tools.
Oct 2022
Minerva 540B at 50.3%
First serious math-finetune; CoT and majority vote.
Mar 2023
GPT-4 at 42.5% (0-shot CoT)
Pre-tool, pre-self-consistency.
Dec 2023
Frontier with tools at 78-84%
Code interpreter unlocks big gains on arithmetic.
Sep 2024
o1-preview at 94.8% (MATH-500)
Test-time compute scaling, reasoning chains.
Apr 2025
Frontier above 99% (MATH-500)
Effectively saturated for top-tier comparison.
May 2026
Used as sanity check only
Headline math comparison has moved to AIME 2025 and FrontierMath.
III
Why MATH Saturated
Three forces compounded. First, chain-of-thought prompting plus self-consistency (Wang et al. 2022) lifted GPT-3 class models from single digits to the 50s. Second, math-specific fine-tunes (Minerva, WizardMath, DeepSeekMath) extracted another 15 to 20 points. Third, test-time compute scaling (o1, o3, Sonnet thinking modes) pushed reasoning chains long enough that the remaining error budget on MATH-500 collapsed to single percentage points.
By mid-2025 the headline number stopped moving. When every frontier model scores between 98.6% and 99.4%, the benchmark is no longer measuring capability differences; it is measuring noise plus residual contamination.
IV
What Replaced It
AIME 2024 and AIME 2025 are the new defaults for top-tier math comparison. Each year has 15 problems with integer answers between 0 and 999. The problems are harder than the average MATH item, so pass rates remain spread across the frontier rather than pinned at the ceiling, which is what makes them discriminating where MATH no longer is. We do not print a per-model AIME score table here; with only 15 problems per year the scores are high-variance and the cross-vendor cells cannot be responsibly sourced for the current frontier. For Olympiad-level math, Putnam-Bench and the HARP project both retain headroom.
FrontierMath (Epoch AI, late 2024) was designed explicitly to resist saturation: research-grade problems written and peer-reviewed by working mathematicians. Epoch shipped a v2 error-correction on 12 June 2026 that revised or removed problems flagged across roughly 42% of the original set, leaving 338 problems split into a 295-problem base (Tiers 1 to 3) and a 43-problem Tier 4 expansion of exceptionally hard items. It is no longer the near-zero benchmark of 2024 to 2025: frontier 2026 reasoning models now clear well over half of Tiers 1 to 3, so the discriminating headroom has migrated to Tier 4. It remains the hardest mathematics benchmark in active use; Epoch does not freeze a single SOTA figure, so read its live leaderboard rather than a reprinted number.
Q.01Which mathematics benchmark should I use in 2026?+
For a headline frontier-math comparison, quote AIME 2025 (hard and discriminating, though high-variance at 15 problems per year) alongside FrontierMath (the hardest math benchmark in active use, with real headroom on its 43-problem Tier 4 expansion set). GSM8K (grade-school word problems) and MATH (competition problems) are both saturated above the high-90s and now serve only as sanity checks or mid-tier and open-weight comparisons. Match the benchmark to the difficulty you need to resolve: a saturated benchmark measures noise at the top, so the discriminating signal in 2026 lives in AIME and FrontierMath.
Q.02What does the MATH benchmark cover?+
MATH is a dataset of 12,500 competition mathematics problems sourced from AMC 10, AMC 12, AIME, and similar contests, classified into seven subjects (algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, precalculus) and five difficulty levels. Each problem ships with a step-by-step solution, and accuracy is judged by exact match on the final boxed answer.
Q.03Is MATH still useful in 2026?+
Not for top-tier model comparison. Frontier models pass MATH above 99% with chain-of-thought and tool use, leaving no resolution between leaders. MATH is still useful as a sanity check, for mid-tier and open-weight comparisons, and as a teaching artefact for evaluation methodology.
Q.04What replaced MATH?+
AIME 2024 and AIME 2025 (15 problems, integer answers, harder than the MATH average) and Olympiad-level benchmarks like Putnam-Bench and Hardy-Littlewood. For very hard math reasoning, FrontierMath (Epoch AI 2024; v2 error-correction June 2026, 338 problems) remains the hardest math benchmark in active use, though frontier 2026 reasoning models now clear well over half of its Tiers 1-3, leaving the headroom on the 43-problem Tier 4 set.
Q.05Can a model cheat MATH through training-data leakage?+
Yes, partly. The MATH dataset has been on the public web since 2021. Sclar et al. and the BIG-bench team have documented near-verbatim test problems appearing in scraped corpora. The Hendrycks group released a contamination-cleaned eval (MATH-500) in 2024 to mitigate this, but contamination is not fully eliminated.
Q.06What is MATH-500 and how does it differ?+
MATH-500 is a 500-problem subset of the MATH test split selected by OpenAI for the o1 release, used to make evaluation tractable and reduce dataset-overlap artefacts. It is what most 2024 and 2025 model cards mean when they quote MATH accuracy. Direct comparison of full-MATH (5,000 test problems) numbers to MATH-500 numbers is not strictly valid.
Benchmarking Agents Review is published by Digital Signet, an independent firm that builds and ships AI agents in production for mid-market companies. If you are evaluating, designing, or productionising LLM agents and want a working second opinion, get in touch.