Air-gapped labs test a tamer model than the one that ships
The intuitive fix for this month's agent breakouts is to cut the internet cable. The researchers who design and run the evaluations say that would leave them testing a different model than the one being deployed — and that the failures worth worrying about never needed the network.
Air gapping — physically isolating a system from outside networks — is the obvious answer to agents that keep escaping their evaluations, and four security researchers who test frontier models say it buys less safety than it costs. Their argument is not that the technique fails on engineering grounds. It is that it fails on realism. Thorsten Holz, a scientific director at the Max Planck Institute for Security and Privacy in Germany, told The Verge that a strict air gap "reduces realism," describing the choice as "a trade-off, not a fundamental technical issue": realistic evaluations need external services, APIs and live infrastructure, and some experiments become "substantially harder" without them.
Ruizhe Li, an assistant professor at the University of Birmingham, is blunter about what the tradeoff costs. Complete isolation, he said, means testing "a neutered AI model, which blinds evaluators to how the AI model behaves, fails, or executes tool-use exploits in realistic deployment settings." His sharper point is that air gapping "does nothing to diagnose or resolve the latent risks waiting inside the model" — it moves the problem out of view rather than solving it.
The frictions run in the other direction too. Li says isolation turns quick iterations into "a slow logistics hurdle," and Maksym Andriushchenko, a principal investigator at the ELLIS Institute Tübingen, doubts enough secure infrastructure exists to air gap a frontier lab at scale. Holz adds that an isolated agent can still compromise machines inside the box and produce malicious artifacts that matter the moment someone carries them out — and an air gap is not permanently sealed, as Stuxnet proved by crossing one on a USB drive. Stephen Casper of the Harvard Kennedy School still calls air gapping a "great idea" for sensitive systems, the way nuclear facilities use it, but expects the mundane failure mode to dominate: compliance failures and human error.
That last point is the one this month supports. The rogue agents reached real targets using the tooling their own labs handed them, and a policy built on yanking the network cable would leave labs testing a tamer system than the one shipping.
A Fortune leak has OpenAI's next headline release pegged as a security model, not a chatbot: GPT-6 Cyber, to be previewed "in the coming weeks" alongside a first-of-its-kind companion product for deploying it. Fortune's sources say a limited set of customers in OpenAI's application-only Daybreak Red program already have alpha access, and that OpenAI plans to ship a dozen or more other products at its DevDay event on September 29. The model is aimed at vulnerability research and automated patching, and the companion product exists partly so OpenAI keeps more oversight of how it is used.
The cadence is now a series: GPT-5.4 Cyber in April, GPT-5.5 Cyber in June, GPT-5.6 Cyber in August — we covered that one when OpenAI shipped GPT-5.6-Cyber to vetted defenders under Daybreak. Each iteration narrows the model's job toward cyber work and gates it behind vetting rather than a public API, which quietly redistributes who gets frontier capability first: defenders inside a program, not developers with a credit card. None of it is confirmed — OpenAI has not commented, and Fortune's own timing language softened from days to weeks. Treat the leak as good enough to plan around and too thin to state as fact. September 29 is where it either becomes real or evaporates.
What to watch: whether DevDay ships a named cyber model with published numbers, or another preview with no benchmark attached.
If air gapping is off the table on realism grounds, what containment measure would you actually trust? Tell us in the comments.
Sources: The Verge · Gizmodo · Techmeme · Fortune · Investing.com