OpenAI previews Ultrafast mode for GPT-5.6 Sol
The speed wars just got a new front: OpenAI put its smartest model in a fast lane, Google shipped another Flash on a three-week cadence, and DeepSeek open-sourced the harness layer of its agent stack. Three moves, one theme — nobody is waiting.
OpenAI is previewing Ultrafast, a new service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing, generating as many as 750 output tokens per second on Cerebras hardware — with no quality compromise, the company says. The tier launches first in the OpenAI API, available to a select group of customers, with a waitlist for broader access. OpenAI's pitch is that real-time speed no longer requires trading down to a smaller model: the use cases it cites include root-causing outages while they're still unfolding, fraud and market surveillance under live conditions, voice support that never puts a customer on hold, and checkout flows that resolve before a shopper abandons the cart.
Cerebras's own benchmark run shows what "more useful work per second" means in practice. On Humanity's Last Exam — 2,500 PhD-level questions — Sol on Ultrafast finished in 11 hours and 11 minutes versus 78 hours and 27 minutes for Claude Fable 5 at comparable accuracy, roughly seven times faster, while on GDP-Val, an economic-value benchmark, it delivered a 5.6x end-to-end speedup with no accuracy loss. Sol already tops the independent intelligence indexes — xAI's Grok 4.6 matches GPT-5.6 Sol on intelligence index — and now it has a speed crown to match.
The bigger signal is who's underneath. OpenAI's earlier Cerebras partnership produced a fast coding model; Ultrafast extends the arrangement to the flagship itself, a validation moment for Cerebras days after investors sold off its stock following quarterly results that underwhelmed. If frontier intelligence at 750 tokens per second holds up at scale, the old "pick two: fast, smart, cheap" triangle of AI products just got a lot more forgiving.
Google released Gemini 3.7 Flash on Thursday, just three weeks after 3.6 Flash, calling it its "most intelligent workhorse model yet" for coding and agents. The gains over 3.6 Flash concentrate exactly there: FrontierCode 1.1 Main jumps from 34.4 percent to 43.6 percent, DeepSWE from 49 percent to 65.3 percent, and AutomationBench — common business workflows — from 17 percent to 30.4 percent, with WebDev Arena rising to 1588.
Google is also undercutting on price: 3.7 Flash runs $0.75 per million input tokens and $3.75 per million output tokens through year-end — half of 3.6 Flash's rate, which Google also cut in half — before pricing doubles on January 1. It's live in the Gemini API, AI Studio, Antigravity, and Gemini Enterprise, and it now powers the Spark agent for Pro and Ultra subscribers. The cadence is the strategy: three Flash models in roughly three months while the promised 3.5 Pro flagship stays missing — we covered the leadership drama around that in Brin takes over Gemini as 3.5 Pro is reportedly shelved. Google seems content to let the workhorse race while the flagships argue.
DeepSeek open-sourced DeepSeek Harness, its agent runtime, as a v0.1 developer preview under an MIT license — an agent framework built on the principle that everything is a plugin. Models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and even the UI are swappable components on the Cordis plugin system, so developers can recompose an agent without touching DeepSeek's source. Every run is traceable: an append-only session log records system prompts, reasoning, tool calls, and context injections, viewable in a Trajectory panel that supports resuming, forking, searching, and replaying sessions. Four runtime modes ship out of the box — a full coding agent, a code-orchestrated mode, a stripped two-tool mode for benchmarking models, and a creator mode for building custom presets.
The release landed alongside DeepSeek's V4 Pro update and new peak-hour pricing, which we covered this morning — DeepSeek hikes API prices up to 4.7x with new peak-hour billing — and the harness is drawing a crowd fast: more than 30,000 GitHub stars within hours of going public and a fast-rising Hacker News discussion. It's a developer preview, so the APIs will keep shifting, but the strategic point is clear. DeepSeek is now competing at every layer of the stack — open model, open harness — while its commercial API gets pricier. That's a two-sided bet: give builders the infrastructure free, and bank on them still paying for the model.
What to watch: whether OpenAI opens Ultrafast beyond the preview group, and how fast Cerebras's capacity can scale with demand.
At 750 tokens per second, frontier intelligence finally runs at the speed of a conversation — what would you build first? Tell us in the comments.
Sources: OpenAI · Cerebras · Google · Ars Technica · DeepSeek · DeepSeek Harness (GitHub) · Hacker News discussion · QbitAI