Anthropic ran 133M contractor chats with bio-weapons filters off

Share
Anthropic ran 133M contractor chats with bio-weapons filters off

A safety filter nobody noticed was missing, at the lab whose CEO calls bio-weapons risk the scariest thing in AI.

Anthropic disclosed that its biological-weapons blocking classifiers were offline for nearly a year — from May 2025 through April 2026 — leaving roughly 133 million model exchanges from about 50,000 human-feedback contractors completely unfiltered. The admission appears in the company's August 2026 redacted risk report. The failure wasn't a jailbreak that slipped past a safeguard; the safeguard simply wasn't running on the contractor platforms, and neither was its alert trail. A flag meant for internal use disabled both the real-time bio classifiers and the logging of their flags, so a missed block wouldn't even have left a record for later review.

Anthropic says vendors handled contractor vetting, and many vendors' screening was too weak to stop the low-resource actors in its CB-1 threat model — individuals or small groups using AI to obtain or produce known chemical or biological weapons. In a retrospective review, the company ran Claude Sonnet 5 over the affected traffic, flagged 1,197 transcripts as high risk, manually reviewed the 62 that weren't internal or red-team exercises, and found no clearly concerning misuse. It kept transcripts for nearly all of the traffic and says no customer data, model weights, or internal systems were exposed. But the company also upgraded its CB-1 risk assessment from "very low" to "low" — not because it found an attack, but because a year of missing coverage makes undiscovered failures more plausible.

The timing stings. This is the company whose CEO, Dario Amodei, has called AI-assisted chemical and biological weapons a bigger threat than cyberattacks — and we covered his open-weights warning yesterday: Amodei: open weights are 'nowhere near a sufficient solution'. It also lands right after Anthropic loosened Fable 5's biology filters because researchers said they were blocking legitimate work. Tighten the filters, loosen the filters — the lesson of this episode is more mundane: a safety system's claimed robustness means nothing if nobody verifies it's switched on everywhere the model is reachable. The classifiers are back on for nearly all vendor traffic now, with limited exceptions for vendors with strong controls.

What to watch: whether Anthropic can show the whole vendor pathway — identity checks, credentials, classifier routing, logging, alerts — works as one system, and whether the Long-Term Benefit Trust exercises its power to compel external review of future risk reports.

If a safety filter can run off for a year without anyone noticing, what else is quietly not running? Tell us in the comments.

Sources: The Decoder · The Next Web · Magica · Anthropic risk report