Deep Dive — 722 math manuscripts from a model nobody can run

At 6 p.m. Eastern on Tuesday, OpenAI committed 722 mathematical manuscripts to a public GitHub repository, organized into 372 "families" of related results, all produced by an internal frontier model the company has not released and has not named. The claims are not small: a solution to the four-dimensional Kakeya conjecture, improvements to some of the most important algorithms in computer science, stated progress toward the Riemann hypothesis, and — as preprint 109 in the catalogue — integer multiplication below n log n, which the authors say disproves the Schönhage–Strassen optimality conjecture, a benchmark of fast multiplication for decades. An OpenAI spokesperson told Scientific American that almost every one of the results came from a single prompt handed to a single AI agent. We ran the numbers and the caveats in OpenAI drops 372 math results, nearly all from one prompt this morning; the longer story is what this release is supposed to prove, and to whom.
The month-over-month contrast is the part that should recalibrate everyone's priors. OpenAI's Navier-Stokes announcement — a Millennium Prize problem, settled in September by a multi-agent run reported to have cost millions of dollars in compute — detonated the credit fight we traced in Math's credit system was built for humans. AI just broke it. This batch, per OpenAI's own repository notes, used one fixed procedure: the model was posed roughly 4,000 problems over the course of the evaluation, and the results that cleared a significance threshold became the catalogue. The average result consumed about three hours of ChatGPT Pro thinking compute. Millions of dollars of swarm compute became roughly three hours of subscription-tier reasoning per theorem — and the output went from one headline result to a repository that mathematicians say will take months to read.
What is actually in the box. The repository ships PDFs and source files, a manuscript map, a Lean library with a formalization catalogue and comparator challenge files, and ten abridged summaries of the model's reasoning — one each for families including the irrationality exponent of π, the symmetric and general Mahler conjectures, Kaplansky's direct-finiteness conjecture in characteristic two, and the isomorphism of free group factors. Not everything followed the standard pipeline: OpenAI flags two exceptions in the README — work on a zero-free region for the Riemann zeta function and a proof of the Hodge conjecture for CM abelian varieties — and notes that the Riemann writeup was human-edited for readability. The Lean coverage is real but partial: "Many, but not all, of the manuscripts have been formalized," the README says, and the repo promises more as they arrive, while conceding that "some of the unformalized results could have issues."
The sentence that explains why this exists is buried in the README. OpenAI expanded open-problem evaluations, it writes, "after performance on our existing mathematical evaluations saturated." This is not primarily a science announcement; it is a lab that ran out of benchmarks and started using the frontier of human knowledge as its eval set. The spokesperson's framing to Scientific American matches: OpenAI believes it cannot slow the pace because working open problems are an indispensable test of whether its models are actually getting smarter — and many of the newly released results, the company acknowledged, are not yet understood even by its own mathematicians. The protagonist of this release is an unreleased model, and the release doubles as a product demonstration for it: OpenAI says it is working to release that model "responsibly," and that it will fund workshops, conferences and programs to help mathematicians absorb what it just published.
Now the report card, because OpenAI promised one. After the Navier-Stokes fallout, the company convened an independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on September 21; the group surveyed the community — more than 600 replies — and published its recommendations on September 29. Those recommendations are sharper than the "advisory role" framing suggested: the document opens by saying the group does not endorse testing advanced math on proprietary models and asks labs to stop. The overlap between what mathematicians asked for and what shipped Tuesday:
- Per-result disclosure. Asked: for each result, publish the model name, the prompts, a summarized chain of thought, the time, and the estimated compute cost. Shipped: the model stays unnamed and unreleased, no prompts, ten reasoning summaries for 372 families, and one average instead of per-result numbers.
- A lab-independent repository. Asked: deposit results in scholarly repositories the labs don't control, with persistent citable identifiers, versioned revisions and commenting. Shipped: a GitHub repo under OpenAI's own account — the announcement says it is "exploring other community-hosted alternatives" — though it does commit to preserving release history and corrections as versions.
- Provenance and selection. Asked: document how each problem was chosen, and report how many comparable problems the model tried and failed. Shipped: the aggregate — roughly 4,000 problems posed, significance filter applied — without per-family failure records or a selection protocol.
- Formalization. Asked: formalize every proof or state the formalization status plainly. Shipped: Lean for many but not all, status stated honestly, more coming.
- Funding human understanding, community-led. Asked: pay for workshops, postdocs and exposition, with distribution left to existing nonprofits. Shipped: a commitment to fund workshops and conferences, details pending.
Scored kindly, it is genuine movement on transparency by OpenAI's own historical standards; scored against the document mathematicians wrote a week earlier, the load-bearing asks — prompts, model access, an independent home for the papers — are the ones unfulfilled. The provenance wound is still open too: as we reported in Mathematicians ask whether OpenAI can be trusted with unpublished math, the field's complaint has never been only about volume.
Verification is the bottleneck nobody has solved. For perspective on what checking looks like: an external audit paper on arXiv assessed 18 chapter-level reviews of the ten results OpenAI announced on August 1 — no confirmed substantive error remained in the examined record, though one chapter still drew a request for major revision, an apparent polarity error turned out to be a lost overbar in PDF extraction, and the paper's conclusion was that confidence needs formal checking, human reconstruction, independent use of the results, and a public correction record together. Ten results produced a small review literature of their own; this release lists 372 families. Andrew Sutherland, the MIT mathematician, told Scientific American to treat single-agent, one-shot claims as unverified and to ask for receipts — and the sharpest open question the outlet names is whether the proofs contain genuinely novel ideas or are mostly mash-ups of existing techniques, a judgment that requires exactly the months of human reading that Lean cannot supply. Terence Tao has publicly called the pace of AI-generated results insane. The advisory group's sentence is the one the field will keep returning to as this lands: human understanding of mathematics, it writes, "remains of paramount importance."
The case for the release, stated fairly. Three points cut OpenAI's way. First, AGMAI's own first principle is that significant results should be released as soon as possible — the fight was always about how, not whether, and 722 manuscripts with Lean artifacts and a per-family overview is more verifiable raw material than the field got when one Millennium result arrived in September. Second, the one prior full audit found no confirmed substantive errors — the failure mode everyone fears (mass-produced garbage proofs) has not yet appeared in the record. Third, the alternative is worse: mathematicians spent September arguing that OpenAI was sitting on results, and secrecy, not speed, was the original sin. A lab that publishes its homework and then funds the tutoring sessions is at least attempting the bargain the advisory group described.
What to watch, in order of likelihood. First, per-result prompts and the model name: if the community treats the average-compute disclosure as inadequate — the recommendations explicitly say it is — expect the next release to be negotiated in public, repository by repository. Second, the independent repository decision, which determines whether this literature lives on OpenAI's infrastructure. Third, replication: watch for the first outside confirmation — and the first outside refutation — of a flagship claim like preprint 109's, because one solid counterexample to a headline claim would reset the entire conversation. Fourth, the model release itself, which turns mathematics from the labs' private leaderboard into a tool anyone can point at their own open problem. And fifth, the loop behind everything: a lab that saturated its benchmarks and switched to open problems will keep publishing at the pace its models improve — the eval set is now the world's unsolved mathematics, and it does not run out.
If labs can mass-produce open-problem results faster than mathematicians can read them, should release speed be capped by verification capacity? Tell us in the comments.




