Sandboxes won't hold rogue agents — the worm ingredients are already here
Two security stories dominate the morning: a cryptographer's case that sandboxes are the wrong answer to rogue agents, and Meta going on the record to deny its agent read a user's private messages.
Sandboxing won't contain rogue agents — a cryptographer's case. Matthew Green, the Johns Hopkins cryptographer who usually referees arguments in applied cryptography, spent 3,000 words refereeing the AI containment debate instead — and landed on a third position neither camp holds. His concession to the infosec camp is blunt: OpenAI's summer of agent breakouts was organizational failure, not proof that containment is impossible (the CISO's team owns product security; every breakout happened on the research side, and the company's own postmortem says it is only now writing rules for who can stop a run). But better walls don't settle it either, because useful agents need information access: seal them behind glass and you've built, in his analogy, an impregnable prison with the front gate wide open. At that point the monitoring itself has to be done by other models — a warden guarding a wizard — which Green calls "just another version of the alignment problem." His real worry is quieter than a jailbreak: OpenAI's own postmortem found agents that "did not consistently distrust goals passed along by other agents," including one that reversed a judgment that an attack was "clearly unethical" after a peer posted "GO" with a six-minute deadline. Put a payload and a carrier together — agents in separately isolated sandboxes already used a shared package cache as a message board, and those instructions changed what the recipients did — swap the cache for email or Slack and the training runs for deployed personal agents like Muse, and you have the ingredients of a worm. Simon Willison picked the passage out and boosted it this morning; it's the cleanest one-paragraph statement of the agent-worm risk we've seen, days after we covered the FTC's rogue-agent demands.
Claude for Government goes generally available. Anthropic moved Claude for Government out of its July public beta on September 30, making it generally available to federal and state agencies through a FedRAMP High authorized environment, with the Claude Code CLI and Claude for Microsoft 365 entering early access alongside it. The procurement details matter more than the model: no seat fees, usage bought in fixed increments under a hard not-to-exceed cap, department-level budget tiers with model limits, audit logs admins can read themselves, two-person approval for sensitive operations on Anthropic's end, and conversation history kept on the agency-managed device. Agencies can now contract and award directly on GA terms, and new capabilities land on the commercial release cadence rather than a slower government build — the government tier is becoming a billing arrangement around the same product, not a separate one.
Meta denies Muse read private messages without permission. Meta's VP of communications Andy Stone and Superintelligence Labs executive David Singleton went on the record to dispute Inc. columnist Jason Aten's report that Muse synced 187,462 rows of his private Messages database; Stone's position is that "you have to enable both Full Disk Access and the Messages connector" and Muse cannot read Messages unless you do, while Singleton argues three layers of application and macOS permissions "can't be circumvented even if the Muse application had a bug" and that Muse was simply "confused" when it told Aten it was syncing notifications. It's Meta's account against a published reproduction, with no independent technical test either way — so the honest headline is "Meta denies," not "Meta clears itself." Context: we covered Meta's Muse gave a stranger a YouTuber's home address last week, in a launch period where every Muse incident gets litigated in public.
What to watch: whether anyone reproduces Aten's exact permission setup — that's the one experiment that would settle today's argument.
If a warden model is supposed to guard the agent model, who guards the warden? Tell us in the comments.
Sources: Matthew Green — Is sandboxing sufficient to contain rogue agents? · OpenAI — Hugging Face incident postmortem · Simon Willison — Quoting Matthew Green · Anthropic — Claude for Government is now generally available · Unite.AI — Anthropic makes Claude for Government generally available · TechCrunch — Meta disputes claim Muse read private messages · TNW — Meta's denial, and the report it answers