Mathematicians ask whether OpenAI can be trusted with unpublished math

Share
Mathematicians ask whether OpenAI can be trusted with unpublished math

The Navier-Stokes credit fight has produced a second, quieter accusation — and it may matter more than the first.

Mathematician Andreas Thom says OpenAI gave him a misleading answer about whether his private work fed its models. After OpenAI announced a result on the non-sofic group problem — work built on methods from Thom and Gábor Kun — Thom emailed OpenAI researchers asking two explicit questions: whether his months of conversations with ChatGPT about the problem had entered training data, and whether they were accessible to the solving process. The reply he got, which he published this week: "Regarding your conversations with ChatGPT: that did not happen." Thom argues the categorical answer actually addressed only direct access during solving, not training — "I take this as dishonesty, to say the least."

Why it matters: OpenAI's own position in the parallel Buckmaster-Alpöge case makes Thom's reading plausible. The company says no specific user data was accessed there, but concedes it "cannot rule out that de-identified data derived from their usage of our products helped improve our models" — which is functionally an admission that researchers' chats can shape future models. Publishing a proof used to take years of quiet, verifiable credit; when a lab's model can ingest your draft thinking and ship the result before you do, the norms of priority that hold mathematics together are suddenly negotiable. And the trust deficit is compounding: the New York Times reports mathematician Tristan Buckmaster is now accusing OpenAI of aggressive behavior in the Navier-Stokes dispute, and OpenAI separately claims progress on yet another Millennium Prize problem.

The uncomfortable part is that "we can't rule it out" is both honest and useless. Mathematicians can't un-share their explorations with ChatGPT, and OpenAI can't retroactively prove a negative about its training mix. Unless the lab publishes a concrete, auditable policy on how researcher conversations are segregated from training data, the rational move for any working mathematician is to stop using frontier chatbots on open problems — which hurts both sides.

What to watch: whether OpenAI responds to Thom's specific charge with evidence rather than a lawyer-shaped non-denial.

Would you feed your best unpublished work to a chatbot whose maker admits it can't rule out training on it? Tell us in the comments.

Sources: Andreas Thom (Mathstodon) · New York Times via Techmeme · The Verge · Science