Princeton's Mengdi Wang: LLMs haven't made a real scientific discovery yet
The models that cracked competition math still can't find a new particle — and Princeton's Mengdi Wang thinks she knows why.
Large models learn "the most likely" answer, and science lives in the long tail. That was the core of a keynote from Mengdi Wang, director of Princeton's Center for AI Innovation, at the Inclusion·Bund Conference in Shanghai on September 10: the training method behind today's LLMs has conquered math and coding but, in her words, still hasn't produced a genuinely new discovery in the fundamental sciences. Her reasoning is structural, not sentimental. Training optimizes toward the maximum of a probability distribution, so models overestimate common cases and underweight rare ones — but in physics, chemistry and biology, the frontier findings are exactly the rare ones, like spotting a star that was never in the training data.
The contrast with math and code comes down to verifiers. A compiler plus unit tests checks every program an agent writes, and formal proof systems like Lean check every step of a proof, so models can try, fail, and iterate against ground truth. Experimental science has no equivalent validator: results depend on personal judgment, specific instruments, and multi-team coordination — one survey in a Nature-family journal found roughly 70% of experiments fail to replicate, and about half the time even the original scientist can't reproduce their own result. Without a verifier, Wang argues, a model just keeps overfitting the known distribution instead of reaching past it.
She also pushed the argument somewhere darker: AI may be slowing science down, not speeding it up. Papers have grown exponentially since LLMs arrived, but the underlying observations of the physical world haven't increased — so the effective signal per paper is falling, and low signal-to-noise makes the real findings harder for anyone to find. Her fix is infrastructure, not model scale: her team is wrapping an automated graphene lab into callable APIs and building LabOS, an "operating system" for research labs where multimodal AI connects models, robots and real instruments, so every experiment is recorded, verifiable and replicable. The bet is that likelihood becomes possibility only when the physical world closes the loop.
We covered the flip side of this earlier this week — OpenAI's Astra cuts the bounded prime gap record to 186 — and the two stories bracket the same gap: verified, formal domains are where AI shines; the unverified physical world is where it stalls.
What to watch: whether verifier-backed setups like LabOS turn "AI for science" from paper generator into discovery engine — the first lab to make its experiments API-checkable wins the argument.
If AI slows science down by drowning signal in paper volume, what should a lab actually do differently? Tell us in the comments.
Sources: Science Net (科学网) · AIbase · Sina Finance (新浪财经)