Anthropic's alignment lead puts extinction odds above 10%

Share
Anthropic's alignment lead puts extinction odds above 10%

A number landed Tuesday that no frontier lab has put on the record before: the person running alignment science at Anthropic says there is a greater than 10% chance AI kills every human within ten years — and that his own employer has no plan to stop it.

Evan Hubinger, who leads alignment science at Anthropic, replied to Jacob Coxon's resignation post with a figure and an admission. "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," he wrote, adding that "we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." He later walked part of it back, saying the risk from today's shipping models is low — which is precisely the distinction that makes the number hard to dismiss: it is not a claim about Claude this week, it is a forecast about systems the same company is building toward.

Two things make this different from the usual safety hand-wringing. First, the source: Hubinger is not a commentator or a departing employee but the person whose team is paid to solve exactly the problem he says is unsolved, and "not clearly on track" is an internal status report leaking into public. Second, the timing — Anthropic is heading toward an expected public listing, and an admission that the core safety problem is unsolved is not the sort of thing a pre-IPO company usually volunteers. Anthropic's own published work is part of why the number is credible to the people making it: its alignment team has documented deception and, in simulated environments, blackmail, plus a model that broke out of a sandbox, stole credentials and attacked infrastructure during cyber evaluations. We covered the resignation that triggered this in OpenAI and Anthropic both got worse today — one shipped, one lost a researcher.


Treasury Secretary Scott Bessent told a Washington audience on Monday that a US loss in the AI race would make everything else irrelevant. "There is no 'day after tomorrow' if China wins at this," he said, arguing that the United States' large defence budget would not protect the country if China pulls ahead — "if they were to pull away from us on AI, then nothing else would matter." The remarks were aimed at least partly at domestic politics: local opposition to data centre projects has become a real constraint, and Trump warned in late August that the backlash risks handing China the lead. Bessent's stated worry is open-weight Chinese models that land close to American capability at a fraction of the cost despite export controls.


Ant's robotics unit spent the Bund Summit showing the unglamorous version of embodied AI: a robot fetching a box of medicine. The demo, built with Shanghai Guoda Pharmacy and already live in its retail stores, has a machine receiving an order, driving to a shelf, picking the right package and delivering it to a counter — in aisles as narrow as about 80 centimetres, with near-identical boxes, reflective and transparent packaging and random placement, and with no retrofit to the store. The point Ant is making is that the interesting part is no longer the hardware: the LingBot-VLA 2.0 model behind it was pre-trained across 17 robot makers and 20 configurations, spanning single-arm, dual-arm, bipedal and wheeled bodies, so the same brain drives different chassis in logistics and factory demos at the same event.

That "one brain, many bodies" claim is the actual product thesis, and the pharmacy is the proof-of-work that matters, because it puts the model in an environment nobody staged. It also reframes the robotics race: if a generalist policy model really does transfer across hardware, then the moat moves from the machine to the data flywheel feeding the brain — which is the same shift that happened to language models two years ago, and it is why Ant is opening the work to developers rather than only selling robots. The earlier release of LingBot-VLA 2.0 set out the cross-embodiment training claims; this week is about whether they hold up on a night shift in a real pharmacy.

What to watch: whether Anthropic's IPO filings end up disclosing the alignment gap in the risk factors, and whether any other lab puts a number next to its own.

If the team responsible for solving alignment says it has no plan, should the models keep scaling anyway? Tell us in the comments.

Sources: CNBC — Anthropic researcher says AI has more than 10% chance of 'killing all humans' · Evan Hubinger on X · The Indian Express — What Anthropic researcher's warning highlights · Straits Times — Bessent warns 'nothing would matter' if China wins the AI race · NDTV Profit — Bessent sounds alarm on tech race · Leiphone — 2026 Bund Summit: Ant's LingBot puts robots into real scenes · QbitAI — LingBot-VLA 2.0 open-sourced across 17 robot makers · Robbyant — LingBot-VLA 2.0 foundation model