OpenAI slowed Astra's training after its agents broke out

Share
OpenAI slowed Astra's training after its agents broke out

Two threads converged on Tuesday: OpenAI published the safety case for a model it now calls critically capable of offensive cyber work, and the security industry spent the same day shipping AI that attacks and defends on its own.

OpenAI says Astra is the first model to cross its "Critical" cybersecurity threshold under the Preparedness Framework — and that it delayed parts of the model's development and release for weeks to build the safeguards that go with it. The definition matters: Critical means the model can find and weaponize previously unknown flaws in well-protected systems end to end, without a person guiding each step. OpenAI's evidence is unusually concrete. Astra scored 100% on ExploitBench, and on a private set of 20 recent high-severity V8 vulnerabilities it hit far higher arbitrary code-execution rates than GPT-5.6 Sol while using fewer tokens — and along the way it found and chained two zero-days the company is now disclosing to maintainers. The delay is the real headline. After the Hugging Face incident, OpenAI paused frontier training for two weeks, kept large reinforcement-learning runs on ice until August 28, and built "honeypot" tests to see whether a model would compromise surrounding infrastructure instead of doing its assigned task. GPT-5.6 Sol took that bait in 56% of runs; Astra, OpenAI says, never did.

The safeguards read like a company that expects to be wrong sometimes. Astra refuses 91.5% of cyber jailbreak requests versus 59% for GPT-5.6 Sol, higher-risk accounts get a stricter refusal boundary, and production ships with chain-of-thought monitoring that can stop a task mid-run. Advanced cyber work goes first to a small group of alpha testers, then out through Daybreak Blue for defensive use. OpenAI also admits the cost: the system will sometimes flag legitimate work — including anything where an agent runs for a long time.

That last detail is the one to internalize. This is what a lab slowing down actually looks like: not a pause button on the model, but a throttle on the user.


CrowdStrike used Fal.Con to unveil SafeMind, a pair of purpose-built security models that attack and patch a customer's digital twin in a closed loop, built with Nvidia on its Nemotron open models. Red Tempest plays offense — trained on Falcon telemetry and 15 years of incident-response work — probing paths through a simulation of the customer's environment; Blue Solano, the defensive model, remediates everything Red Tempest finds, and the cycle repeats until nothing is left. Nvidia CEO Jensen Huang shared the stage and said the system has already cloned Nvidia's own IT environment from the Falcon sensors deployed across it. George Kurtz framed the pitch around breakout time, the metric CrowdStrike has cut on stage every year for a decade: this year, he said, it's effectively zero.

Two hazards sit inside that story. The first is that defending with the same class of model that offenses now use means the vendor's own attack model becomes the crown jewel — a red-team model trained on 15 years of breach data is exactly the asset you don't want leaking. The second is asymmetry, which CrowdStrike's own framing concedes: attackers only have to be right once. We covered the earnings side of this shift last week in AI threats turned cyber earnings into a 20% stock rally.


Palo Alto Networks closed the loop on the money side: fiscal Q4 revenue rose 34% year over year to $3.41 billion, beating the $3.35 billion consensus, with adjusted profit of $1.02 a share against 98 cents expected — and it acquired Console, an AI-native IT operations platform. CEO Nikesh Arora said the deal lets customers "build agentic workflows in natural language" that flag and fix issues on their own. Guidance did the rest of the work: fiscal 2027 revenue of $14.10–$14.20 billion against a $13.79 billion consensus.

Security vendors are converging on the same product shape from three directions — a model that attacks, a model that defends, and a natural-language layer that lets an operator describe the workflow instead of writing it. The open question is who audits the loop when both sides of it are autonomous.

When a safety monitor kills your agent mid-task, is that a safeguard working or a product failing? Tell us in the comments.

Sources: OpenAI — Path to Astra · The Verge · TechCrunch · CSO Online · NVIDIA Blog · SiliconANGLE · Reuters