Sarnak's AlphaZero test — LLMs can calculate, not create math
The people best placed to judge whether AI can do mathematics keep landing on the same verdict — and one of them just proposed a clean test for when the machines would finally prove him wrong.
Princeton number theorist Peter Sarnak has proposed an "AlphaZero Test" for AI mathematics — and he expects the machines to fail it. In the August issue of Notices of the American Mathematical Society, Sarnak — the Eugene Higgins Professor at Princeton and an emeritus professor at the Institute for Advanced Study — lays out the challenge: hand a theorem-proving machine the Ramanujan Conjecture in the elementary form Srinivas Ramanujan originally posed it, give it no database and no prior theory, and ask it to decide whether the statement is true. Sarnak postulates the machine will come back stumped, because the proofs behind most of mathematics' deepest theorems lean on abstraction, conceptualization, and bodies of accumulated theory — precisely what statistical learning, starting from zero, does not reconstruct.
Sarnak, a semi-professional chess player before he became a mathematician, borrows the frame from DeepMind's AlphaZero, the self-taught engine that rewrote chess and Go by passing what he calls a "statistical complexity threshold" special to those games. Mathematics is a different universe, he argues: it is infinite, its rules are given, and Gödel's incompleteness guarantees statements that cannot be decided. He expects the threshold AlphaZero cleared for chess to hold for math — and if he is wrong, he says, the impact would be dramatic, because we would hold theorems proven true without understanding why, forcing a rethink of what mathematics even is. Either way, he sees theorem provers as a Stockfish for mathematicians: powerful assistants that rapidly derive what he calls the "elementary statistical derivatives" of existing theories — the problems most working mathematicians spend most of their time on.
Sarnak's skepticism is converging with the field's. Fields Medalist Timothy Gowers argued days ago that LLMs are strong at combining known methods and searching many paths but lack the intuition to pick the few productive routes in a vast search space — we covered his essay this morning in Gowers: LLMs solve math's biggest problems with counterexamples. DeepMind researcher Tom Zahavy's ICML position paper "LLMs Can't Jump" names the missing ingredient "manipulative abduction": inventing new foundational assumptions with no linguistic precedent, something he suggests world models might eventually supply.
The through-line across all three is not that AI is useless at math — it is that the division of labor is getting clearer. Machines can derive, search, and check at a scale no human can match, and Sarnak expects them to become indispensable assistants. The bottleneck is creativity: choosing the fruitful direction, inventing the abstraction nobody saw coming — exactly where each of these observers, independently, places the limit. The essay is refreshingly honest, down to Sarnak's admission that many mathematicians have their "heads in the sand, ignoring AI, hoping it goes away." And it was his former student Jacob Tsimerman who pushed him to the point: what would AI have to do before Sarnak took it seriously?
What to watch: whether any lab takes up the AlphaZero Test as a benchmark — a from-scratch proof of a deep theorem, trained on nothing but the statement.
If an AI proved a deep theorem no human could verify line by line, would you accept the result? Tell us in the comments.
Sources: AMS Notices — Sarnak's "AlphaZero Test" · The Decoder · Zahavy, "LLMs Can't Jump" (OpenReview) · Gowers's Weblog