Deep Dive — Tier 4 fell in 14 months. Nothing can refill it
The hardest closed-form math benchmark in the industry is over, and the interesting fact is not the winning score. It is that every candidate replacement is structurally worse at the one job a benchmark has to do: letting someone outside a lab check whether the lab is telling the truth.
As this morning's brief reported, Epoch AI now treats FrontierMath Tier 4 as saturated. GPT-6 Astra cleared the final unsolved item — a research-level question written by combinatorialist Jay Pantone — taking the top score to 97.6%, with GPT-5.6 Sol at 83.0% and GPT-5.6 Terra at 68.3%. AlphaSignal notes the tier launched in mid-2025 with the best model of the day clearing 5% of it. Fourteen months, wall to floor.
That is a short news item with a long tail, because Tier 4 was not just another leaderboard. It was the field's explicit answer to "your benchmarks are all saturated" — a purpose-built instrument designed to stay hard for a decade. Its collapse in barely more than a year says something about the pace, and something else entirely about how we measure capability at all. The second thing is the one nobody has settled.

What Tier 4 actually was
Epoch describes FrontierMath as hundreds of original, exceptionally hard problems written and vetted by expert mathematicians, with a typical Tier 4 item requiring multiple hours of a researcher's time and the upper end multiple days. The current dataset holds 338 problems: a base set of 295 across Tiers 1–3, plus 43 in Tier 4. Twelve are public. Solving a problem means producing a closed-form answer that a grader can check mechanically — no proof essay, no human referee, no taste.
One number on Epoch's own page deserves more attention than it has received. On June 12, 2026, Epoch shipped a major revision "addressing errors in 42% of problems." Nearly half the items in the hardest general-purpose reasoning benchmark the field has were flawed enough to need correction, and that correction landed a year after Tier 4's release. The instrument that was supposed to be more rigorous than scraped web data turned out to need its own audit — which is the honest baseline for everything below.
There is also the relationship, disclosed by Epoch in one sentence: FrontierMath "was developed with funding from OpenAI, who has exclusive access to a subset of the benchmark." OpenAI paid to build a test and gets to see part of it that other labs cannot. That arrangement predates the saturation result and was contentious at launch. It becomes more awkward now that the sponsor's model has won the whole thing.
The replacements are harder to trust, not just harder to solve
Epoch has already moved ground twice. FrontierMath Erdős is a set of 68 conjectures posed or studied by Paul Erdős, formalized in Lean 4, curated by Thomas Bloom, who maintains erdosproblems.com and estimates that only three to five Erdős problems of this caliber had previously fallen to AI. Fifty of the 68 Lean statements come from Formal Conjectures, Google DeepMind's open library of formalized open problems; the rest Epoch autoformalized with Bloom reviewing each one. Grading is strict and genuinely mechanical: submissions must pass Comparator, a checker maintained by the Lean FRO, compiled in a sandboxed container with no network access, accepted only if the submission proves the identical statement using permitted axioms. Astra solved 2 of 68 in its official run, rising to 5 with repeated attempts that AlphaSignal reports cost more than $220,000 in compute.
Read those two numbers together and you get the real picture of machine mathematics in September 2026: a model that closes a research-grade answer-only benchmark at 97.6% manages under 8% when the task is "write a complete, machine-checkable proof in a formal system." Tier 4 asked for the destination. Erdős asks for the road, and the road is where models still fail.
The problem is that the road scales badly as a measurement device. FrontierMath took 338 problems from more than seventy volunteer authors and grades answers mechanically. The Erdős set needed a specialist curator, a formalization pipeline, human review of every autoformalized statement, and a dedicated proof checker. That is a fraction of the throughput at an order of magnitude the cost, and there are not many mathematicians with Lean fluency and a list of open problems they're willing to spend weeks formalizing. The same scarcity applies to the third arm, Epoch's Open Problems track, which collects significant unsolved research questions whose solutions happen to be computationally verifiable — a set that is, by construction, limited to whatever unsolved problems exist that also have checkable answers.
The dependency chain inside the replacement is worth following. The Erdős set has one curator deciding which problems are hard enough to count. Fifty of its 68 Lean statements were lifted from a Google DeepMind library; the other seventeen needed an autoformalization pipeline plus human review item by item. A benchmark now rests on one person's judgment about difficulty, on another company's formalization of what the problems mean, and on a checker maintained by the Lean standards body. Each link is defensible. Together they mean the number of institutions whose cooperation a capability claim requires has gone from "some mathematicians" to "a specific small set of people and their tools."
The number is becoming a compute figure
The subtler breakage is that benchmark scores have stopped meaning "capability" and started meaning "capability times budget."
Look at the cleanest result of the week. Nvidia published the full recipe behind Nemotron's IMO 2026 run — 30 of 42 points, above the gold-medal cutoff of 29 — and alongside the score released the entire compute ledger: roughly 707 million generated tokens and 1,464 GB200 GPU-hours, with three post-trained checkpoints generating 384 candidate proofs per problem and a proof accepted only when sixteen independent verifier judgments unanimously scored it perfect. OpenAI's own Tier 4 run was, by its description, maximum reasoning effort. AlphaSignal flags the consequence directly: that setting inflates both latency and token spend.
Once the sampling budget is a free variable, a benchmark score is not one number but a curve, and different labs report different points on it. Two teams can both "solve Tier 4" at a fifty-fold cost difference and both be telling the truth. And the verifiers themselves are part of the system being graded: Nemotron's model-based jury estimated 32 points where the official marking gave 30, with the whole overestimate concentrated on two problems — the judges shared a blind spot rather than noising independently. A benchmark whose scoring depends on models from the same generation it is evaluating is measuring something, but it is no longer measuring it from outside.
A February paper, "When AI Benchmarks Plateau," put numbers on the general pattern: across 60 language-model benchmarks scored on 14 saturation-related properties, nearly half showed saturation, with the rate climbing as benchmarks aged. The one design feature reliably associated with staying hard was expert curation — not keeping the test set private. Which is precisely the resource that is expensive, slow, and, in FrontierMath's case, partially funded by the company at the top of the table.
What the mathematicians are objecting to is a different thing
This is where the measurement argument and the argument Terence Tao and 24 fellow Fields Medalists made on Friday collide, and the collision is worth stating because almost nobody has noticed it.
The declaration Tao posted — signed by medallists from Deligne in 1978 through Yu Deng this year — does not claim AI is bad at math. It claims that "solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight," that famous problems function as "landmarks and lighthouses" for a body of technique, and that "the mass production at faster and faster pace of 'true/false' statements could destroy fertile ground instead of breathing life into new ideas." Tao wrote that the group deliberately skipped the slower consultative process behind the earlier Leiden declaration because the situation felt urgent.
The evaluation community's problem is the mirror image. It needs a continuous supply of new hard, checkable questions and is short of exactly that. The mathematicians' complaint is that those questions were never the point — they were seeds for technique, and the seeds are being spent as fast as possible.
Both complaints are correct and they cannot be satisfied by the same pipeline. A recent paper from a group working on scalable mathematical discovery makes the trade-off concrete by describing the supply chain: the FAR system starts from 5,245 combinatorics papers, extracts 6,453 candidate conjectures, filters to 4,717 that appear well-posed and still open, surfaces 598 potential resolutions and hands 77 to a human author team to review. The authors' framing is blunt — frontier-model reasoning is the scarce resource and expert human review is "even more sharply constrained," so the bottleneck in AI mathematics is no longer solving. It is choosing what to solve and then checking.
Automate problem generation and you get abundant, cheap, verifiable items — and a benchmark that measures nothing the mathematicians care about, since a machine-composed conjecture has no community waiting behind it to turn the solution into method. Don't automate it and you cannot produce hard test items at anything like the rate the models consume them. Tier 4's 43 problems lasted fourteen months. There is no bench of volunteer experts that can refill that annually.
We already ran the credit side of this argument when OpenAI's Navier–Stokes claim landed, and the deep dive on why mathematics' priority system broke still stands: verification, not generation, is the binding constraint, and nobody funds it. What the last week adds is that the same constraint now applies to evaluation itself. The people who can write a hard research-grade problem and the people who can check a proof are the same small population, they are unpaid, and the labs are hiring them.
The contrarian read: the crisis is useful
Give the skeptics their case, because there is a strong one.
First, "the benchmark is saturated" and "mathematical reasoning is solved" are unrelated claims, and the Erdős number is the proof: 2 of 68. Anyone concluding from Tier 4's death that research mathematics is finished is reading the wrong column. Second, saturation is the normal fate of every fixed test, and the field has escaped worse traps — GLUE, ImageNet, SWE-bench's early splits all aged out and the work continued. A dead benchmark is an inconvenience for marketing decks, not for science. Third, the labs benefit from an evaluation crisis: if the only credible tests are private, curated, expensive and gated behind a handful of institutions with conflict-of-interest statements on their own websites, then capability claims become unfalsifiable in a very convenient direction. The mathematicians' demand for slower, more transparent announcements and the industry's move toward closed, high-cost evals point the same way — toward fewer people being able to check.
Where that read breaks is on the substance. Verification capacity has not grown at all, the price per test item has gone up rather than down, and the one open, auditable artifact this week came from Nvidia shipping a reproducible recipe with its compute bill attached. Openness is currently the scarce good in AI mathematics, not compute.
What to watch
Three things will tell you whether this is a rerun of past benchmark deaths or a permanent change.
Whether any Tier 4 replacement gets published as an open set. FrontierMath Erdős has released its Lean statements and its evaluation scaffold, which is the single strongest trust signal available, and it should be checked whether the same holds for whatever comes after Tier 4 — including whether the funder still gets exclusive access to part of it.
Whether problem supply gets automated and how fast scores decay once it does. The FAR pipeline is a proof of concept that conjecture extraction at scale is possible; the first benchmark built mostly from generated problems will be the real test of whether hardness and meaningfulness can be produced separately.
Whether the labs start reporting cost alongside capability. A number without a token budget attached is no longer comparable across systems, and nobody has proposed a standard. That is a solvable engineering problem that nobody has solved because solving it hurts the leaderboard.
The uncomfortable summary: fourteen months was the half-life of the hardest evaluation anyone could design with unpaid experts and closed-form answers, and every successor trades one form of checkability for another. We are not running out of ways to test models. We are running out of ways to test them that outsiders can verify.
Should frontier benchmarks be funded by a public body rather than the labs they rank? Tell us in the comments.
Sources: Epoch AI — FrontierMath Tier 4 · Epoch AI — FrontierMath Erdős · AlphaSignal · Terence Tao — A Severe Misalignment of AI in Mathematics · arXiv — When AI Benchmarks Plateau · arXiv — The Problem Is the Problem