FrontierMath: 338 Research-Grade Problems, Four Tiers
The hardest mathematics benchmark in active use. Original problems written and peer-reviewed by working mathematicians, most taking a domain researcher hours and the upper tier days. Models solved under 2 percent at launch; the headroom that remains now lives almost entirely in Tier 4.
What FrontierMath measures
FrontierMath was introduced by Epoch AI in November 2024 (Glazer, Erdil, Besiroglu and colleagues) to answer a specific problem: the standard mathematics benchmarks had saturated. GSM8K and MATH were both scoring above the high-90s for frontier models, which means they no longer separate one strong model from another. FrontierMath was built to sit far above that ceiling.
The problems are original and unpublished, authored and peer-reviewed by expert mathematicians, and they span most major branches of modern mathematics: computationally intensive questions in number theory and real analysis through to abstract questions in algebraic geometry and category theory. Solving a typical problem requires multiple hours of effort from a researcher in the relevant branch, and the upper-tier questions take multiple days. At publication, the paper reported that state-of-the-art AI models solved under 2 percent of the problems. That near-zero floor is what earned FrontierMath its reputation as the hardest math benchmark in active use.
The tier structure and the v2 error-correction
The current dataset is 338 problems, split into two sets. The v2 error-correction, released on 12 June 2026, revised or removed problems that had been flagged for errors or ambiguity in the original v1 set and is the version to quote today.
The tier split matters for interpretation. Tiers 1-3 run from undergraduate difficulty up to advanced-graduate exploratory problems; Tier 4 is genuine research-level mathematics, with some problems expected to resist AI solution for a long time. Because the tiers differ so sharply, a single overall FrontierMath percentage hides most of the signal. A model can look strong by clearing a large share of Tiers 1-3 while barely touching Tier 4, and the Tier 4 number is the one that actually tracks the research frontier. Quote the tier set, not just an aggregate.
Scoring and contamination defence
FrontierMath uses automated verification of exact answers rather than an LLM judge. The problems are designed so that a correct solution produces a specific, checkable answer, which keeps scoring reproducible and cheap and removes the judge-model variance that weakens softer benchmarks. Crucially, the problems are unpublished: they have never been released in plaintext, so a model cannot have memorised the answers from a training crawl in the way that has been demonstrated for MMLU and HumanEval. See our contamination explainer for how memorisation distorts the older benchmarks.
Unpublished problems are a structural contamination defence, not merely an empirical one. That is the same design principle that makes GAIA's private test set trustworthy. It is the reason a hard, held-out mathematics benchmark can stay meaningful even as the models that attack it are trained on ever-larger slices of the public web.
The OpenAI funding and access question
FrontierMath is also a case study in why benchmark provenance matters. On 20 December 2024, around the launch of OpenAI's o3 model, Epoch AI disclosed that OpenAI had funded the creation of FrontierMath and had access to the problem statements and solutions for the bulk of the set. Many contributing mathematicians had not been told of OpenAI's involvement, and the late disclosure drew criticism (TechCrunch, January 2025). Epoch published its own account of the arrangement and the safeguards.
The safeguard that matters for reading scores is the held-out set: a group of problems for which AI developers receive only the statements, not the solutions, so that at least part of the benchmark tests models on material no lab has had prior access to. The practical rule for a reader is data-provenance, not suspicion of a particular figure: a score reported on problems a lab had access to is not directly comparable to a score on the held-out set, so read any self-reported frontier number alongside which portion of the benchmark it used. FrontierMath remains a valuable instrument; the access history is simply part of what a careful reader keeps in view.
Reading the scores in 2026
We do not reprint a per-model FrontierMath score table. The benchmark has moved well off its near-zero 2024 floor: frontier 2026 reasoning models now solve a substantial share of Tiers 1-3, so the discriminating headroom has migrated to Tier 4, which remains largely unsolved. Epoch AI does not freeze a single headline figure; it maintains a live leaderboard that separates the tier sets and the dataset versions and moves continuously as new models are evaluated.
So the honest way to cite FrontierMath is a triple: the model, the tier set (Tiers 1-3 or Tier 4), and the dataset version (v1 or v2). A bare "X percent on FrontierMath" with none of those attached is not a comparable number. To choose the right benchmark for a use case rather than chase a frontier headline, start from the homepage task picker.
When to use FrontierMath
FrontierMath is the right instrument when you need to resolve differences at the very top of mathematical reasoning, the regime where MATH and GSM8K have saturated and can only measure noise. It is the natural companion to Humanity's Last Exam as a saturation-resistant frontier benchmark, and it pairs well with AIME 2025 for a fuller top-tier picture.
For most application and product evaluation it is the wrong tool: the problems are research-grade, and a model can be excellent for your workload while scoring near zero on Tier 4. If you are ranking models for a real deployment, evaluate on your own data first and treat FrontierMath as a capability-ceiling signal, not a product-fit test. For everyday reasoning comparison use the reasoning benchmarks guide instead.
Q.01What is FrontierMath?+
Q.02How many problems does FrontierMath have and what are the tiers?+
Q.03What was the v2 error-correction?+
Q.04Why does the OpenAI funding matter for reading FrontierMath scores?+
Q.05What is the current best score on FrontierMath?+
Q.06Should I use FrontierMath to evaluate a model?+
Sources
- [1] Glazer, E., Erdil, E., Besiroglu, T. et al. (2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. arXiv:2411.04872.
- [2] FrontierMath overview, Epoch AI. epoch.ai/frontiermath.
- [3] FrontierMath Tiers 1-3 (v2), released 12 June 2026. epoch.ai/benchmarks/frontiermath-tiers-1-3-v2.
- [4] FrontierMath Tier 4 (v2), 43-problem expansion set. epoch.ai/benchmarks/frontiermath-tier-4-v2.
- [5] Clarifying the creation and use of the FrontierMath benchmark, Epoch AI. epoch.ai/latest/openai-and-frontiermath.
- [6] AI benchmarking organization criticized for waiting to disclose funding from OpenAI, TechCrunch, 19 January 2025. techcrunch.com.