Independent reference. No vendor affiliation. Reprinted scores dated to their source; elsewhere we link the official board. Editorial policy.
Abstract
What338 unpublished, expert-vetted problems; 295 in Tiers 1-3, 43 in Tier 4 (v2).
WhoGlazer, Erdil, Besiroglu et al., Epoch AI, 2024 (arXiv:2411.04872).
ScoringAutomated verification of exact answers; problems held privately, never published.
Leaderboardepoch.ai/benchmarks
Section I.v Knowledge and Reasoning|Reviewed 2026

FrontierMath: 338 Research-Grade Problems, Four Tiers

The hardest mathematics benchmark in active use. Original problems written and peer-reviewed by working mathematicians, most taking a domain researcher hours and the upper tier days. Models solved under 2 percent at launch; the headroom that remains now lives almost entirely in Tier 4.

Which frontier-math benchmark should I quote in 2026?Quote FrontierMath for the very top of mathematical reasoning, where MATH and GSM8K are saturated above the high-90s and measure only noise. Pair it with AIME 2025 for a fuller top-tier picture: AIME is discriminating but high-variance at 15 problems per year, while FrontierMath is research-grade and its Tier 4 set is still largely unsolved. Always name the tier set and dataset version (v1 or v2) with any score.
I

What FrontierMath measures

FrontierMath was introduced by Epoch AI in November 2024 (Glazer, Erdil, Besiroglu and colleagues) to answer a specific problem: the standard mathematics benchmarks had saturated. GSM8K and MATH were both scoring above the high-90s for frontier models, which means they no longer separate one strong model from another. FrontierMath was built to sit far above that ceiling.

The problems are original and unpublished, authored and peer-reviewed by expert mathematicians, and they span most major branches of modern mathematics: computationally intensive questions in number theory and real analysis through to abstract questions in algebraic geometry and category theory. Solving a typical problem requires multiple hours of effort from a researcher in the relevant branch, and the upper-tier questions take multiple days. At publication, the paper reported that state-of-the-art AI models solved under 2 percent of the problems. That near-zero floor is what earned FrontierMath its reputation as the hardest math benchmark in active use.

II

The tier structure and the v2 error-correction

The current dataset is 338 problems, split into two sets. The v2 error-correction, released on 12 June 2026, revised or removed problems that had been flagged for errors or ambiguity in the original v1 set and is the version to quote today.

Set
Problems
What it tests
Tiers 1-3
295
The base set. Difficulty ranges from undergraduate problems up to exploratory problems suitable for an advanced graduate student. This is where most measurable movement has happened as frontier reasoning models improved through 2025 and 2026.
Tier 4
43
The expansion set of exceptionally hard, research-level problems, crafted as short research projects by professors and postdocs. Designed to surpass Tier 3 in difficulty; the discriminating headroom in 2026 lives here.

The tier split matters for interpretation. Tiers 1-3 run from undergraduate difficulty up to advanced-graduate exploratory problems; Tier 4 is genuine research-level mathematics, with some problems expected to resist AI solution for a long time. Because the tiers differ so sharply, a single overall FrontierMath percentage hides most of the signal. A model can look strong by clearing a large share of Tiers 1-3 while barely touching Tier 4, and the Tier 4 number is the one that actually tracks the research frontier. Quote the tier set, not just an aggregate.

III

Scoring and contamination defence

FrontierMath uses automated verification of exact answers rather than an LLM judge. The problems are designed so that a correct solution produces a specific, checkable answer, which keeps scoring reproducible and cheap and removes the judge-model variance that weakens softer benchmarks. Crucially, the problems are unpublished: they have never been released in plaintext, so a model cannot have memorised the answers from a training crawl in the way that has been demonstrated for MMLU and HumanEval. See our contamination explainer for how memorisation distorts the older benchmarks.

Unpublished problems are a structural contamination defence, not merely an empirical one. That is the same design principle that makes GAIA's private test set trustworthy. It is the reason a hard, held-out mathematics benchmark can stay meaningful even as the models that attack it are trained on ever-larger slices of the public web.

IV

The OpenAI funding and access question

FrontierMath is also a case study in why benchmark provenance matters. On 20 December 2024, around the launch of OpenAI's o3 model, Epoch AI disclosed that OpenAI had funded the creation of FrontierMath and had access to the problem statements and solutions for the bulk of the set. Many contributing mathematicians had not been told of OpenAI's involvement, and the late disclosure drew criticism (TechCrunch, January 2025). Epoch published its own account of the arrangement and the safeguards.

The safeguard that matters for reading scores is the held-out set: a group of problems for which AI developers receive only the statements, not the solutions, so that at least part of the benchmark tests models on material no lab has had prior access to. The practical rule for a reader is data-provenance, not suspicion of a particular figure: a score reported on problems a lab had access to is not directly comparable to a score on the held-out set, so read any self-reported frontier number alongside which portion of the benchmark it used. FrontierMath remains a valuable instrument; the access history is simply part of what a careful reader keeps in view.

V

Reading the scores in 2026

We do not reprint a per-model FrontierMath score table. The benchmark has moved well off its near-zero 2024 floor: frontier 2026 reasoning models now solve a substantial share of Tiers 1-3, so the discriminating headroom has migrated to Tier 4, which remains largely unsolved. Epoch AI does not freeze a single headline figure; it maintains a live leaderboard that separates the tier sets and the dataset versions and moves continuously as new models are evaluated.

So the honest way to cite FrontierMath is a triple: the model, the tier set (Tiers 1-3 or Tier 4), and the dataset version (v1 or v2). A bare "X percent on FrontierMath" with none of those attached is not a comparable number. To choose the right benchmark for a use case rather than chase a frontier headline, start from the homepage task picker.

VI

When to use FrontierMath

FrontierMath is the right instrument when you need to resolve differences at the very top of mathematical reasoning, the regime where MATH and GSM8K have saturated and can only measure noise. It is the natural companion to Humanity's Last Exam as a saturation-resistant frontier benchmark, and it pairs well with AIME 2025 for a fuller top-tier picture.

For most application and product evaluation it is the wrong tool: the problems are research-grade, and a model can be excellent for your workload while scoring near zero on Tier 4. If you are ranking models for a real deployment, evaluate on your own data first and treat FrontierMath as a capability-ceiling signal, not a product-fit test. For everyday reasoning comparison use the reasoning benchmarks guide instead.

Editor's verdictFrontierMath is the hardest mathematics benchmark in active use and the best current instrument for resolving frontier-lab math capability. Quote the v2 dataset, name the tier set, and read Tier 4 as the live research frontier. Keep the OpenAI access history in view when comparing self-reported numbers, and read Epoch's live board rather than a reprinted figure.
Reader Questions
Q.01What is FrontierMath?+
FrontierMath is a mathematics benchmark from Epoch AI, introduced in November 2024. It is a set of original, unpublished, exceptionally difficult problems authored and peer-reviewed by expert mathematicians, spanning most major branches of modern mathematics from number theory and real analysis to algebraic geometry and category theory. Solving a typical problem takes a domain researcher multiple hours, and the hardest problems take multiple days. At publication, state-of-the-art AI models solved under 2 percent of the problems, which is why FrontierMath is described as the hardest math benchmark in active use.
Q.02How many problems does FrontierMath have and what are the tiers?+
After the v2 error-correction released on 12 June 2026, the full dataset is 338 problems. It splits into a 295-problem base set called Tiers 1-3, which ranges from undergraduate-level difficulty up to problems suitable for an advanced graduate student, and a 43-problem expansion set called Tier 4, which is genuine research-level mathematics. Tier 4 problems are crafted as short research projects by professors and postdocs; some are expected to resist AI solution for a long time. Because the tiers differ so sharply in difficulty, the tier breakdown is more informative than any single overall number.
Q.03What was the v2 error-correction?+
FrontierMath v1 (late 2024) contained problems that were later found to have errors or ambiguities. On 12 June 2026 Epoch AI released a v2 dataset that revised or removed the flagged problems, leaving 338 clean problems (295 in Tiers 1-3 and 43 in Tier 4). This is a normal part of maintaining a hard benchmark: as strong models attempt the problems, disputed items surface and get corrected. When quoting a FrontierMath score, name the version (v1 or v2) and the tier set, because they are not directly comparable.
Q.04Why does the OpenAI funding matter for reading FrontierMath scores?+
Epoch AI disclosed on 20 December 2024 that OpenAI had funded the creation of FrontierMath and had access to the problem statements and solutions for the bulk of the set. Epoch also maintains a held-out set of problems whose solutions no AI developer receives, which is the intended contamination safeguard. The practical reading rule: a score from a lab that had access to problems and solutions is not directly comparable to a score on the held-out set, and self-reported frontier numbers should be read alongside which portion of the benchmark was used and who had prior access. This is a data-provenance question, not a claim that any specific number is wrong.
Q.05What is the current best score on FrontierMath?+
We do not reprint a frozen per-model score here. FrontierMath has moved well off its near-zero 2024 floor: frontier 2026 reasoning models now solve a substantial share of Tiers 1-3, while Tier 4 remains largely unsolved and is where the discriminating headroom now lives. Epoch AI publishes a live leaderboard that moves continuously and separates the tier sets and dataset versions, so read the current figures there rather than a reprinted number, and always quote the tier set and version alongside any score.
Q.06Should I use FrontierMath to evaluate a model?+
Use FrontierMath when you need to resolve capability differences at the very top of mathematical reasoning, where saturated benchmarks like MATH and GSM8K measure only noise. For most product and application evaluation it is overkill: the problems are research-grade and a model can be excellent for your use case while scoring near zero on Tier 4. FrontierMath is a frontier-lab and research-comparison instrument, best quoted alongside AIME 2025 for a fuller picture of top-tier math ability.
MATH BenchmarkReasoning Benchmarks ComparedHumanity's Last ExamGPQA-Diamond and ARC-AGIBenchmark ContaminationWhat Benchmarks Miss

Sources

  1. [1] Glazer, E., Erdil, E., Besiroglu, T. et al. (2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. arXiv:2411.04872.
  2. [2] FrontierMath overview, Epoch AI. epoch.ai/frontiermath.
  3. [3] FrontierMath Tiers 1-3 (v2), released 12 June 2026. epoch.ai/benchmarks/frontiermath-tiers-1-3-v2.
  4. [4] FrontierMath Tier 4 (v2), 43-problem expansion set. epoch.ai/benchmarks/frontiermath-tier-4-v2.
  5. [5] Clarifying the creation and use of the FrontierMath benchmark, Epoch AI. epoch.ai/latest/openai-and-frontiermath.
  6. [6] AI benchmarking organization criticized for waiting to disclose funding from OpenAI, TechCrunch, 19 January 2025. techcrunch.com.
From the editor

Benchmarking Agents Review is published by Digital Signet, an independent firm that builds and ships AI agents in production for mid-market companies. If you are evaluating, designing, or productionising LLM agents and want a working second opinion, get in touch.

Book a 30-min scoping callDigital Signet →

30 minutes, free, independent.·1-page action plan within 48h.·Honest if not the right fit.