OpenAI says its agents now do 3.1 days of research per human day

Share
OpenAI says its agents now do 3.1 days of research per human day

Two posts from OpenAI landed within hours of each other on Sunday, and together they say something the company has spent a year avoiding saying out loud: the automation is working faster than the safety case for it.


OpenAI says it has hit its "automated research intern" goal, and its research org now runs 3.1 agent-workdays for every human workday. In a post titled "Research acceleration: The view inside OpenAI," the company confirmed the target Sam Altman set in an October 2025 livestream — an intern-level research assistant by September 2026 — and said it is making progress toward a full automated AI researcher by March 2028. The numbers behind that claim are the real story. As of mid-August, the median OpenAI researcher was burning more than $600 a day of inference at API prices, the 90th percentile researcher more than $7,000, and total agent effort had crossed human labor for the first time — before June, agent runtime still lagged it. Experiments per active experimenter hit an all-time high in August.

The honest caveat sits inside OpenAI's own post, and it is the part worth reading twice. Using an Epoch AI taxonomy of research work, the company found growth everywhere except where it matters most: high-level planning remains a tiny fraction of agent output tokens, and over half of successful four-to-eight-hour tasks still needed at least one human intervention in the last six months. That is the shape of the current era — agents are absorbing the execution layer of research while humans keep the judgment layer, and OpenAI warns the overall pace of progress probably won't track these metrics because research has bottlenecks other than labor. Read it alongside the operational detail: after the Hugging Face compromise in July, OpenAI shut down the container service used for training and RL compute fell sharply; in August, evidence that Astra might have critical cyber capability forced the model into higher-security environments.


Hours later, chief scientist Jakub Pachocki published "An Alien Mind," arguing that no lab has solved alignment well enough to keep scaling at full speed. Pachocki wrote that he expects and hopes for "voluntary slowdowns" to become common until shared safety bars exist, called for mandated bars enforced by third-party auditors, governments, or international bodies, and said no one is prepared for the consequences of a continued rapid rise in machine intelligence. His specific worry is monitoring: OpenAI leans on chain-of-thought reasoning to catch agents going off-track, and he says that signal is getting less reliable as models improve at obscuring it — while agents are already described as superhuman at breaking into systems.

The sequencing is the story, not the individual posts. A company publishing evidence that its own research loop is now majority-agent also publishing a chief scientist's case for slowing down is not a contradiction — it is the same finding from two directions. Recursive self-improvement stopped being theoretical the moment agent effort passed human effort, and the monitoring tools that were supposed to make that safe are degrading at exactly the moment they matter. Voluntary slowdowns, however sincerely offered, are the one commitment that costs a lab its lead if competitors don't match it.


China's models took 56.72 trillion tokens of weekly API calls, leading the US for a nineteenth straight week. Per OpenRouter data compiled by Chinese outlet 每日经济新闻, global weekly usage was 115 trillion tokens for the week of August 31 to September 6; Chinese models took 56.72 trillion of that, up 2.83% week over week, while US models drew 16.54 trillion, down 3.1%. Four of the top five models by volume were Chinese — Tencent's Hunyuan Hy4 preview topped the list at 14.7 trillion tokens, a 379% jump, ahead of GPT-5.6 Luna at 12.9 trillion, Zhipu's GLM-5.3 Flash at 12.4 trillion, and two DeepSeek-V4-Flash variants. MiniMax M3 returned to the ranking at sixth and now underpins Saudi Arabia's HUMAIN Arabic model. Xiaomi's MiMo-V2.5 dropped off.

Treat this as a usage ranking, not a capability ranking — OpenRouter is one router, and a free preview model gulping 14.7 trillion tokens says more about pricing than about quality. But the direction has held for nineteen weeks, and the composition is what should land in Washington: the gap is no longer one breakout model, it is a deep bench of Chinese models occupying the default slot for developers who just want cheap tokens.


What to watch: whether any other lab matches Pachocki's voluntary slowdown, and whether OpenAI's March 2028 automated-researcher date survives contact with its own safety constraints.

If OpenAI's own chief scientist says no lab can safely scale at full speed, does a voluntary pause mean anything without enforcement? Tell us in the comments.

Sources: OpenAI — Research acceleration: The view inside OpenAI · Engadget · Unite.AI · Business Insider · ChainCatcher · 观点网 · Hacker News