Meta raced to patch a VM escape in Muse before launch

Share
Meta raced to patch a VM escape in Muse before launch

Three stories today share a theme: systems that were supposed to be contained — an agent platform, a preprint archive, a text watermark — all straining at the edges.


Meta's own security teams didn't think Muse was ready to ship. 404 Media reports that in the weeks before launch, engineers found several vulnerabilities in the viral agent product, at least one of which could have let a normal Muse user break out of the sandbox and reach sensitive internal Meta databases. The evidence is an internal post by three infrastructure executives — VP of core infrastructure Surupa Biswas, VP of engineering Francois Richard, and senior director of engineering Josh Barry — sent to the core infrastructure team on September 18, ten days after launch, describing "a sudden spike in reported KVM escapes" that made them "rally on a service hardening push" starting August 27. Muse shipped September 8, about twelve days into that push; the issues reached Mark Zuckerberg, and staff worked nights and weekends. A Meta source described the result as "half-baked protections being rushed out to enable the launch," adding that "many senior engineers believe it's inevitable we're going to have a massive data breach as a result of Hatch" — the internal name for Muse. Meta's own bug-bounty page pays up to $300,000 for a Muse VM escape, its highest-tier payout, precisely because each agent runs in a per-user VM sitting one kernel flaw away from production. Security researcher Patrick Wardle put it plainly: "having access to production environment literally one KVM escape away, is plain irresponsible." We covered the sealed-VM design when the product arrived — Meta ships Muse: a consumer agent inside a sealed VM — and this scoop shows the boundary was thinner than the architecture diagram suggested. Meta said in a statement it has "strengthened Muse through extensive dogfooding, agentic red teaming and our bug bounty program."


arXiv is capping researchers at two papers a month. The preprint archive announced a hard rate limit effective October 1: two submissions per calendar month per submitter, with at most three active at any time — its first across-the-board cap, after a year of escalating fights with AI-generated slop. The arithmetic is hard to argue with: arXiv received 40,363 submissions in September 2026, against 20,569 in September 2024 and 9,869 in September 2016, and the computer science category has grown roughly sixfold in two years. Thomas Dietterich, chair of arXiv's editorial advisory council, said a small fraction of authors flood the queue with low-quality papers "consuming a disproportionate fraction of the moderators' time," delaying everyone else's work by days or weeks. arXiv calls AI tools a direct driver of the flood. We flagged the funding side of this squeeze last month — arXiv just funded its independence — as AI floods its queue — and now the moderation dam is holding by decree. Whether a per-submitter cap survives coordinated abuse is the obvious next question.


OpenAI starts watermarking ChatGPT text — in the EU first. Answering the EU AI Act's machine-readable-provenance requirement, OpenAI opened global API opt-in to text watermarking today and will add invisible watermarks to eligible ChatGPT and Codex output in the European Union over the coming weeks, with detector access restricted to approved researchers and expert organizations rather than the public. The technology, called textGrain, embeds a statistical signal in word choices, and unusually for this kind of announcement, OpenAI published its failure modes along with it: at a 1% false-positive rate the detector catches about 80% of 200-word passages (roughly 95% at 400 words), and replacing 10% of words with synonyms drops detection from about 92% to 66%. That candor is the right call — a watermark presented as proof would be worse than none. What to watch: whether Meta discloses further Muse vulnerabilities as bounty hunters dig in, and how arXiv actually enforces its new cap.

Is a personal agent that sits one kernel bug from production really "contained"? Tell us in the comments.

Read more

Altman says the world must accept AI's 'bad things'

Altman says the world must accept AI's 'bad things'

A heavy news day for AI governance and open weights: OpenAI's CEO is publicly pricing the trade-off his industry keeps dodging, Reflection finally put specs on the model it teased yesterday, and AMD is trying to set the terms before Nvidia's RTX Spark lands. Altman says the world should accept AI's "bad things" — and the labs' new pact agrees. In an interview released Monday on Politico's Decoded podcast, Sam Altman said OpenAI's position is "we believe that the world should accept some bad th

Today in AI — October 5, 2026

Today in AI — October 5, 2026

The day the ecosystem stopped pretending everyone is a partner: Meta and Microsoft quietly cut their Claude budgets, Washington gave AI policy a new name, and New York City put lab executives under oath. Elsewhere, one model learned to drive a robot, and Mac users finally got Apple Intelligence off their disks. Models & Research * Reka AI's Rho-1 collapses the multimodal stack into a single 19-billion-parameter model. The research preview runs text, images, video and robot control as token

OpenAI adds text watermarking to ChatGPT and Codex — EU first

OpenAI adds text watermarking to ChatGPT and Codex — EU first

Regulation is now shipping inside the product: OpenAI's EU-only watermark rollout lands today, Wikimedia publishes its evidence against OpenAI's agents, and two of Anthropic's biggest customers are easing off Claude. OpenAI is turning on invisible text watermarking in ChatGPT and Codex — starting with the European Union. Over the coming weeks, eligible EU users across all plans will get a machine-readable signal called textGrain woven into the text the model produces, while API customers anywh

The Take — A diary in Claude isn't a written threat

The Take — A diary in Claude isn't a written threat

I think charging Carli Michelle Heller with a second-degree felony over a sentence she typed into Claude at 5:10 a.m. stretches Florida's written-threat statute past recognition. A message addressed to nobody is not a writing transmitted "in any manner in which it may be viewed by another person" — unless the only person who views it is your chatbot vendor's safety reviewer, and if that is the rule, nothing you type into any moderated app is private anymore. Our morning brief and yesterday's d