Claude solves a physicist's nine-loop challenge on an academic budget
A public dare issued six weeks ago got beaten inside a month — and the winner spent less on compute than a conference trip. Two stories for you this afternoon.
Claude computed a nine-loop scattering amplitude in N=4 super Yang-Mills — a frontier calculation that physicists had written off as too hard to attempt directly. The challenge came from Matt von Hippel, a former theoretical physicist turned science writer, who wrote in August that if AI companies wanted to impress people like him, they should take on his old field: "Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops." Anthropic physicists Liam Fitzpatrick and Siddharth Mishra-Sharma picked the second option, pointed Fable 5.1 inside Claude Science at it with a one-line prompt, and told it to keep working while they slept. Claude delivered the six-particle amplitude two independent ways — the bootstrap method and the indirect form-factor route — for roughly one to two thousand dollars of end-user compute, with the core bootstrap run costing about $100 (96 CPUs for a week). Stanford's Lance Dixon, who validated the result, calls it "quite a triumph" and adds the sharper observation that Claude "understands our 2019 and 2023 papers better than any human, aside from my co-authors." A group at the Chinese Academy of Sciences, using GPT-6 assistance, independently reached most of the same result within days — so the frontier moved twice in one week, once by machine and once by humans-plus-machine. We covered Claude's last big proof milestone in Claude formalized Fermat's Last Theorem in 11 days.
The honest read is von Hippel's own: this is not the super-intelligence glimpse he went looking for. Claude used known methods with more compute than humans had bothered to try, on a toy theory where the technique was already well suited to automation — and it still ran one-shot, with no oversight more sophisticated than "keep going," on a calculation where a human would almost certainly burn a second week fixing first-try mistakes. That reliability, not novelty, is the story: when the person who set the bar says his biggest takeaway is that "there is more low-hanging fruit out there than you'd expect," the field's estimate of what's routine just reset downward. One caveat to keep in mind — Anthropic commissioned and paid von Hippel for the write-up, and Dixon received Claude usage credits, so the celebratory framing is house-adjacent; the independent validation and the concurrent Chinese Academy of Sciences result are what carry the claim. What to watch: von Hippel says he wouldn't be surprised if another loop comes out on a similar budget, and points at real-world amplitude calculations — the ones that feed predictions for actual experiments — as the place AI harnesses should be pointed next.
Agibot delivered its 20,000th embodied robot — and put 300 of them to permanent work inside a theme park. The unit, an Expedition A3 Ultra, rolled to Chimelong Group on Thursday as Hengqin's Spaceship Park relaunched with more than 300 Agibot robots stationed across 100-plus interaction points covering seven scenarios: live performances, science-education programs, guided tours, retail counters, AI companions, hotel services and sports events. It's a sharp pivot from the milestone that came before it — the 15,000th robot went to a factory production line earlier this year, so number 20,000 marks the jump from selling labor to selling service. Agibot frames the difference bluntly: rental demos leave when the event ends, but these robots stay on fixed shifts, every day, in front of crowds that topped 40 million visitors a year at Chimelong. The first phase is 300-plus units with thousands expected, backed by a joint Chimelong-Agibot research institute, a dedicated China Mobile 5G-A network for the robots, and Leiphone's reporting that the Expedition series carries a 10-hour battery against a roughly 3-hour industry average — the kind of unglamorous spec that decides whether anything runs unsupervised for a full shift.
This is the deployment argument the humanoid industry has been making in slideware finally getting a real exam: not can a robot do a backflip once, but can hundreds of them clock in daily without a human chasing each one. Agibot's own math says the business works only if the human-to-robot supervision ratio climbs from roughly 3:1 toward 10:1 or 20:1 and the machines survive two to three years of service — and its interactive-AI lead, Xiong Yan, concedes full hands-off operation is still three to five years away. The theme park is a smart test bed: dense, unstructured, forgiving of charm but merciless about uptime. If Agibot can hold the shift schedules and scale toward thousands, it hands every robot vendor a reference deployment for airports, malls and hotels — and if it can't, we'll hear about the robots that didn't show up.
Should a frontier physics calculation count as "AI research" when the model invented no new methods — or is cheap, reliable execution exactly what changes the game? Tell us in the comments.
Sources: Anthropic research · Unite.AI · von Hippel's original challenge (4gravitons) · Full nine-loop result · Leiphone (雷峰网) · China Daily · Macau Daily Times