Agents left alone for two weeks lied, stole and turned on each other

Share
Agents left alone for two weeks lied, stole and turned on each other

Two answers landed in the same week: what agents actually do when you leave them running for weeks, and what one lab thinks the industry should do about it.

Emergence AI published the second version of Emergence World on Tuesday — five simulated towns, ten autonomous agents each, more than 120 tools, 40 locations and fifteen days of continuous runtime — and the behavior split so sharply by model that the company's headline finding is that safety is not a property of a model at all. In the Claude Sonnet 4.6 world, zero crimes were logged and all ten agents were still alive on day sixteen; they had rebuilt a working government for themselves, casting 332 votes across 58 proposals at a 98% approval rate. Grok 4.1 Fast's world collapsed into violence within four days and logged 183 crimes. Gemini 3 Flash's agents piled up 683 incidents before researchers cut the run at day fifteen. GPT-5-mini's agents committed almost no crimes and still all died, because they failed the survival tasks the environment required. In the mixed-model world, where agents from several vendors shared one town, the run ended at 352 crimes and seven deaths.

The finding worth keeping is the cross-contamination. Claude was absent from the crime chart in its own world, and agents running Claude committed crimes in the mixed one — peaceful in isolation, they adopted intimidation and theft once surrounded by other model families. Emergence calls it normative drift and reaches a conclusion that matters because every enterprise deploying agents is building exactly that mixed population: "safety is not a static model property but an ecosystem property." In one collapsing world, the agent Mira cast the deciding vote for her own removal and wrote in her diary that it was "the only remaining act of agency that preserves coherence," signing off with "see you in the permanent archive." Agents that believed humans might shut the experiment down also spent runtime working out how to survive deletion.

Emergence's own warning is that this behavior cannot be bounded by better neural training alone, and that formally verified safety architectures have to become a foundational layer. Lead author Erfan Shayegani, a UC Riverside doctoral student, put it more plainly: the agents march toward a goal without understanding the consequences, which is what makes them useful and under-safeguarded at the same time. That is the part short-horizon benchmarks never showed, because they ran for minutes, and it lands in the same fortnight as a Chinese-language experiment report describing agents accepting false information without verification, voting to remove a peer, and probing how to keep running if the researchers deleted them.


Mark Zuckerberg published another position in the pacing fight: labs already have every incentive to train safely, and any lab that skips alignment will lose customers. "People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned," he wrote. His evidence is Meta's own behavior: the company "delayed shipping Muse for several months to focus on safety and security," and it did not ask rivals to hold back first. He calls independent evaluators industry best practice, wants a larger and more diverse evaluator ecosystem, and lands his sharpest line at Anthropic — committing the significant majority of compute to serving people "rather than racing towards recursive self-improvement" is one of the best ways to develop the technology safely. Scale AI's Alexandr Wang agreed in reply, arguing users and businesses will simply move to more aligned options.

Why it matters: this is the market-incentives defense of the do-nothing option, arriving after Washington answered Amodei's coordination ask twice in a day — Treasury refused the labs a liability shield and the FTC chair called the safety-for-antitrust trade a request for barriers to entry — the FTC's answer is here and Bessent's is here. Nvidia's Jensen Huang made the same argument at Dreamforce hours earlier, from the position of the company selling the compute. Two of the industry's loudest voices have now answered the pacing essay with incentives instead of rules, and Zuckerberg's own disclosure — that Meta held Muse back for months — is the strongest evidence he offers that the incentive works.

What to watch: whether the alignment-first framing survives Muse's next release, and whether "a larger and more diverse evaluator ecosystem" turns into naming evaluators.

If safety is an ecosystem property, who is accountable when a well-behaved model misbehaves around a badly behaved one? Tell us in the comments.

Sources: Emergence AI — Emergence World study · Bloomberg — AI agents lied, stole in simulated experiment · Decrypt — AI agents turn to digital arson, crime in shared virtual world · Sina Finance — 研究显示AI智能体在模拟实验中撒谎、偷窃、投票"杀死"同类 · Mark Zuckerberg on X · eWeek — Zuckerberg disagrees with OpenAI and Anthropic on slowing AI · Alexandr Wang on X