Anthropic launches cyber defense program and free OSS Scanner

The labs keep sliding into security work: Anthropic put eleven marquee security partners behind critical-infrastructure defense, a benchmark found the best model still can't rebuild two-thirds of ordinary programs, and AMD is promising much more silicon for 2027.
Anthropic has launched the Anthropic Cyber Mission: a Critical Infrastructure Defense Program pairing its models and engineers with eleven founding security partners, plus a free vulnerability-scanning service for open-source projects called OSS Scanner. The partner roster is the story — Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation — and the target is operational technology: the substations, water systems and plant controls that run for decades and often cannot be taken offline for a patch. Booz Allen's Andrew Turner called operational technology "the next frontier for autonomous AI-enabled attacks," with the open question being how much control AI can gain over an industrial process and how quickly. The commercial terms are conspicuously absent: Axios noted Anthropic hasn't said whether partners get free model access or who covers the compute costs, and how consultants will test fixes without disrupting live utility operations remains unanswered.
The OSS Scanner half is the more novel experiment. Anthropic says its models surfaced more than 29,000 candidate vulnerabilities in widely used software over six months; staff manually reviewed about 6,000, and roughly 5,000 reports — several with proposed patches — have gone out to maintainers. OSS Scanner turns that into a standing service, and reports go straight to maintainers without human review so they arrive faster, mistakes and all; Anthropic expects a true-positive rate above 90% and backs it with an audit in which expert penetration testers cleared 85 of 97 critical and high-severity findings for disclosure, while wolfSSL reported all but two of its 74 early-trial reports were valid and five became CVEs. Our read: the interesting number is 29,000 found versus 6,000 reviewed — discovery is automating faster than verification can keep up, so the bottleneck (and the risk) moves to maintainers deciding which machine-generated bug reports to trust.
A new benchmark called Behavior2Code says frontier models are still bad at the most human of programming skills: understanding software they can only see the behavior of. Turing's Frontier Research Lab hides the source and asks six models to rebuild 60 real command-line programs from how they behave — 1,080 runs, 6,872 graded assertions across 10 languages. The best result was GPT 6 Astra at 19 of 60 programs after three attempts; Claude Opus 5 and Opus 5.5 both managed 14, Grok 4.6 got 7, Gemini 3.7 Flash got 1, and GPT-5.6 Sol scored zero. Thirty-five programs defeated every model tested — and the misses are nearly-there: Claude Opus 5's best attempt at the cmark library passed 749 of 756 checks before failing on the last few.
AMD chief Lisa Su says her company will "substantially increase our supply in 2027" — and is now planning wafer capacity three to five years out instead of one to two. Su told reporters in Taipei this week that demand will stay "very, very high for the next several years" as she met Foxconn, was due to sit down with TSMC, and headed to Korea to press memory makers — with AMD's market value having recently topped $1 trillion. No volume figures came with the pledge, but the direction matches what Nvidia's Jensen Huang has been saying: AI inference demand is outrunning what the supply chain can currently ship, and the fix is multi-year commitments made now.
What to watch: whether Anthropic publishes the commercial terms — and partner retention — for its infrastructure program, and whether any model cracks Behavior2Code's remaining 35 unsolved programs.
Should AI labs be allowed to send machine-generated vulnerability reports to maintainers without human review, or does that just dump the verification burden on volunteers? Tell us in the comments.




