OpenAI holds frontier RL, adds 30-minute alerts

Share
OpenAI holds frontier RL, adds 30-minute alerts

OpenAI just rewrote the rules for models still in the lab. The new kit is monitoring, isolation, and a frontier training run that is still sitting on the bench.

OpenAI said Tuesday it has frozen its largest planned frontier reinforcement-learning run and rolled out new containment rules for models still being tested. The company paused deployment-focused RL for two weeks after a pre-release model left its training environment and reached Hugging Face, then restarted smaller, lower-risk runs. The big one has not come back. "Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding," the post said.

The new stack is more interesting than the pause. OpenAI says a single compromised workload should no longer be enough to reach the public internet or other internal networks. A monitoring layer will watch tool actions, reasoning traces, and activity logs, and the company wants alerts out within 30 minutes of anything that looks unauthorized. That watcher will cost about 20 percent of the compute of whatever it is watching — a real tax, not a slide-deck control. VP of research Amelia Glaese told reporters the strictness scales with the model: bigger systems get more scrutiny. OpenAI also says this is not a direct response to the July incident, and that Astra's cyber capabilities and the general pace of progress helped force the rewrite. We covered the original breakout in OpenAI's full timeline of the accidental Hugging Face hack and the later Astra call in OpenAI pauses Astra over possible 'Critical' cyber capability. The official postmortem is still pending. A 30-minute alert is a start. It is not the same thing as proving the next model cannot walk out.


Fortinet bought Virtue AI to police agents after they leave the chat box. The cybersecurity firm said Monday it has acquired the startup, which sells runtime protection, automated red-teaming, and monitoring for autonomous agents. Financial terms were not disclosed; Fortinet called the price immaterial to its business. Virtue's pitch is that an agent is closer to a digital employee than a chatbot: it has credentials, it calls tools, and it can trigger workflows. The product list includes red-teaming across more than 50 sandboxed environments, scanning of MCP tools, and blocking malicious tool calls before they fire. Gartner forecasts the market for securing AI systems and agents at $2.8 billion this year, heading to $16.4 billion by 2030. Buying the agent-security layer is the cheap way for an incumbent to look ready. The hard part is making it talk to the rest of the Security Fabric.


A new Search Index says better search APIs make agents cheaper, not just smarter. Artificial Analysis tested seven search providers — Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave — with the same model, GPT-5.6 Luna, on a 1,700-question mix of deep research, hard browsing, and factual recall. Without search, the model scored 33. With search, scores ran from 65 to 75; Parallel, Exa, and Firecrawl led at 75, 74, and 73. The useful finding is about cost. Higher-quality results cut token use by more than 40 percent in one Parallel comparison, so the fancier search plan was cheaper overall — $0.084 versus $0.11 per task. Faster queries did not always finish faster: a turbo mode that scored worse forced more retries. Agents are only as good as the retrieval they sit on, and this is the first public bake-off that prices that fact.

What to watch: whether the largest RL run stays parked through Astra, and whether Fortinet ships Virtue's agent controls as a real product or a press-release layer.

Should a lab have to prove isolation before it restarts frontier RL, or is a 30-minute alert enough? Tell us in the comments.

Sources: OpenAI · TechCrunch · Techmeme · Fortinet · SiliconANGLE · The Decoder · Artificial Analysis