The neurosurgery resident who solved Crouzeix's conjecture — with ChatGPT 5.6

Share
The neurosurgery resident who solved Crouzeix's conjecture — with ChatGPT 5.6

A 22-year-old open problem in numerical linear algebra fell this week — not to a tenured professor, but to a neurosurgery resident who let GPT-5.6 Sol run for 16 hours and went back to clinical work while it worked. The mathematician who posed the problem has checked the proof and believes it is correct. Meanwhile, the crypto industry is asking AI labs to hand defenders the same frontier cyber tools attackers already use.

A neurosurgery resident at Peking Union Medical College Hospital has solved Crouzeix's conjecture — a famous open problem in numerical linear algebra — with the key step produced during a roughly 16-hour autonomous run of GPT-5.6 Sol in ChatGPT Work mode, and numerical analysts Alex Townsend and Anne Greenbaum, along with Michel Crouzeix himself, have checked the argument and believe it is correct.

The solver, Dr. Shanmu Jin, is a postdoctoral researcher and neurosurgery resident with an unlikely résumé: an undergraduate degree in geology, then an M.D., with all mathematics beyond standard science courses self-taught. He found his way into matrix analysis through research on transcranial ultrasound, was drawn to the conjecture's deceptively simple statement — for any matrix A and polynomial p, the norm of p(A) is bounded by twice the maximum of p on the numerical range of A — and on July 27 posted a preprint claiming a proof. Townsend, who had spent a year asking ChatGPT to attack the problem and watched it stall at the same missing lemma every time, initially read the preprint with skepticism. Within hours, he and Greenbaum realized the argument was real.

Jin's method matters as much as the result. He adapted the prompt OpenAI used when solving the Cycle Double Cover conjecture: the system was denied web access, instructed to spin up many subagents exploring genuinely different proof strategies without converging prematurely, to subject candidate proofs to adversarial audits, and to keep going until a complete proof survived checking. Jin started the run and did not intervene. The decisive theorem surprised the experts — rather than the heavy estimates the field expected, the argument reduces the problem via a careful sampling strategy to a simple positivity condition. Jin open-sourced everything: the prompt, successive manuscripts, a Lean formalization, and an axiom audit.

The context makes the moment land. Crouzeix posed the conjecture in 2004; he proved a bound of 11.08 in 2007, and with Palencia improved it to 1+√2 (about 2.414) in 2017, the year a dedicated week-long workshop at the American Institute of Mathematics ended with no resolution. The inequality matters practically — it transfers approximation error from the complex plane to bounds on matrix functions, underpinning workhorse numerical methods like GMRES and Krylov solvers. And the proof was not a fluke of one prompt: eight days after Jin's preprint, Emiel Lorist and Felix Schwenninger posted an independent five-page proof — disclosing that they, too, used ChatGPT 5.6 to explore strategies. As Townsend and Greenbaum wrote, "It is remarkable to realize that a longstanding conjecture was first solved by someone with no specialized training in mathematics, working with a large language model."


More than 40 digital-asset organizations — including Coinbase, Block, BitGo and Strategy — signed an open letter asking frontier AI labs to give vetted defenders the same cyber capabilities attackers already wield.

The letter, coordinated by the Bitcoin Policy Institute, asks labs for early access to frontier cyber models, sufficient compute, secure research environments, and direct channels with lab security teams. The argument is one of asymmetry: safety guardrails routinely block legitimate defensive security work — even approved researchers get refused when defensive tasks resemble offensive activity — while attackers face no such limits and can run open or locally deployed models. The demand lands as labs widen access anyway: OpenAI recently split its cyber offering into Daybreak Blue and Red and released GPT-5.6-Cyber for advanced defensive work that would normally trip stronger safeguards. The coalition's point is that security researchers securing the Bitcoin ecosystem should not have to work with the models' safety restrictions still on.

What to watch: whether math's verification pipeline can absorb a flood of AI-produced proofs, and whether the labs' new trusted-access programs actually quiet the defenders' complaint.

If a self-taught clinician can crack a 22-year-old conjecture with a 16-hour model run, does the gatekeeping of mathematical discovery need to change? Tell us in the comments.

Sources: SIAM News · AI Era (Xin Zhi Yuan) · Jin's preprint · Independent proof (arXiv) · CoinDesk · crypto.news · CryptoSlate