Ten Claude agents formalize a 122-year-old physics problem
A proof no human wrote, a 320-billion-parameter release from an unknown Chinese lab, and OpenAI's billing system quietly starting to carry other labs' models — the overnight lane had a research flavor.
Ten Claude Sonnet 5.5 agents spent roughly 15 hours and 1,270 messages producing a 17,895-line Lean proof that settles the N=7 case of the Thomson problem — J.J. Thomson's 1904 question of how seven charged points arrange themselves on a sphere to minimize their mutual repulsion. Vals AI, the benchmarking startup behind the work, says the proof builds cleanly in 599 seconds across 8,928 jobs, depends only on three standard axioms, and was independently accepted by a second, separately built proof kernel — which also rejected a deliberately mutated version of the theorem, the negative control that makes the double check mean something. The formal statement carries real caveats, and the team says so up front: its README notes the statement is "not human-certified" and not peer reviewed — a machine checking a proof only proves the encoded statement, and no mathematician has yet signed off that the encoding is the right problem — and the Thomson problem for general N remains open; the agents closed one case, not a century. That honesty is what separates this from a blog-post claim: the artifact is public and re-checkable, which is exactly the bar we've held for the genre — Claude solves a physicist's nine-loop challenge on an academic budget.
Chinese lab IQuest Research has open-sourced IQuest-Q1, a 320-billion-parameter mixture-of-experts model built for command-line coding agents — about 15 billion active parameters per token, a 512K context window, and weights already downloadable. The release's best story comes from the team's own training pipeline: when a reinforcement-learning reward curve stalled, the model was handed a CLI harness, read its own training logs and traces, and traced the failure to an extra space its inference service was injecting during text decoding — fixing it got the early training rounds moving again. Treat that as a good demo, not a doctrine. The benchmarks, including a claimed 84.5 on CyberGym, are self-reported by a lab few outside China have heard of, and one self-debug on the team's own scaffold is not "models fix themselves now" — but a capable 320B open-weight coding model joining the commons is worth the download regardless.
Kimi K3 can now be billed against an existing OpenAI enterprise commitment — via Baseten, not Moonshot. The inference provider says it is among the first open-model providers in OpenAI's B2B Marketplace, with a native integration inside Codex: enterprise customers can run Moonshot's Kimi K3 — alongside Zhipu's GLM-5.3-Flash and Whisper Large V3 — through Codex or the Responses API and draw the spend down against the same procurement commitments they already have with OpenAI, currently behind an interest-form waitlist. The "first Chinese model inside OpenAI's billing system" framing belongs to Chinese financial media and is shakier than it sounds, since GLM 5.3 is Chinese too; OpenAI didn't announce the deal itself, and no uptake, pricing or revenue figures exist. The durable shift is the plumbing: OpenAI is becoming the front door for other labs' open weights, which is procurement consolidation wearing a partnership's clothes.
What to watch: whether OpenAI co-signs the Baseten marketplace story, and how far Google widens Argon's Fairwind rollout this week.
Would you accept a proof of a 122-year-old problem if no human had read it line by line? Tell us in the comments.
Sources: Vals AI — Thomson N=7 Lean proof · thomson-n7-lean (GitHub) · Startup Fortune · IQuest-Q1 model card (Hugging Face) · Pandaily · QbitAI via iFeng Tech · Baseten — OpenAI partnership · TechNode