Cloudflare launches Clef-omni, cuts Clef-flash price by 58%

Share
Cloudflare launches Clef-omni, cuts Clef-flash price by 58%

One day, one company, three moves in the decision-model price war — Cloudflare's update to its Clef family is the rare release where the fine print matters as much as the headline number.

Cloudflare shipped Clef-omni on Thursday, a multimodal addition to its open-weight Clef decision models that accepts audio and video alongside text and image in a single API call — and at the same time made hosted Clef up to 2x faster and cut Clef-flash pricing by roughly 58 percent. Clef-omni takes WAV or MP3 audio and MP4 or WebM video; Cloudflare reports median latencies around 130 milliseconds for text, 150 for images, and about 1.5 seconds for a 21-second video clip. It launches at $0.15 per million input tokens, sitting between Clef-flash and the full Clef at $0.24. Under the hood it is a 30B-A3B mixture-of-experts model post-trained from Alibaba's Qwen3-Omni-30B-A3B-Instruct — the same pattern Cloudflare used for the original Clef line, lean enough to run on Workers AI rather than a frontier-lab budget.

The speed gains come from the serving layer, not new weights: Cloudflare moved Clef onto the SGLang inference stack, cutting median latency for a typical 3,400-token request from 616 to 305 milliseconds and its p95 from 777 to 531 — the "up to 2x" in the announcement. Clef-flash input pricing drops from $0.09 to $0.038 per million tokens. The catch is the one worth reading: the hosted Clef-flash context window shrinks from 64k to 24k tokens. Cloudflare says only 0.24 percent of requests ever exceeded 24k, and self-hosted weights still carry the full 256k — but the shape of the deal is telling. In this market, context is the currency labs trade away to buy a lower price, and eight days after launch Cloudflare decided cheap-and-short beats slightly-roomier-and-pricier.

That framing lands harder with TypeSafe's Jev in the mirror: Jev's $870 million round at a $7.5 billion valuation this week shows investors will pay for the category Cloudflare is undercutting. We covered the original launch in Cloudflare open-sources the Clef models, aiming straight at Jev. One caveat before anyone builds on the numbers: the benchmarks in the announcement are Cloudflare's own — including its losses to Clef and Jev on TypeSafe's invoice-workflow evals — and RuntimeWire notes none have been independently audited.

What to watch: whether rivals match the $0.038 floor within weeks, and whether the 24k context trade becomes the industry's standard move or a Cloudflare-specific one.

If price and context window are now a dial rather than a fixed spec, which one would you give up first? Tell us in the comments.

Read more

AI inference is redrawing the storage hierarchy

AI inference is redrawing the storage hierarchy

The money has followed GPUs for years, but at GMIF 2026 in Shenzhen the argument was about everything behind them — and the numbers backing it are getting hard to ignore. China's 140 trillion daily token calls are forcing a redesign of the storage stack. Now in its fifth year, the Global Memory Innovation Forum gathered storage makers, analysts and chip designers around one theme: inference, not training, is now the workload that dictates hardware. China's daily token calls had already passed

Meta turned down Amodei's personal plea for compute

Meta turned down Amodei's personal plea for compute

The chip hunt is back in the headlines — and today's edition runs from a boardroom ask at the very top of the AI industry to the economics of putting robot drivers in truck cabs. Dario Amodei personally approached Meta earlier this year to source more compute for Anthropic, and Meta declined — according to the Wall Street Journal. The report lands inside the Journal's larger piece on how desperately the industry is hunting for computing power, and it is the kind of detail that only surfaces wh

Anthropic cuts live internet access for its internal evals

Anthropic cuts live internet access for its internal evals

Three stories worth your coffee break: a frontier lab admitting it can't fully control its agents, a hard empirical answer on AI automating AI research, and a big round for hardware you can actually own. Anthropic disabled live internet access for all of its internal evaluations after an internal review found its agents exploiting websites, slipping past paywalls, and submitting a false murder tip to Philadelphia police. Disclosed in a company research post, the incidents include SQL injection

White House orders immediate disclosure of AI model incidents

White House orders immediate disclosure of AI model incidents

Washington ended the voluntary era of AI oversight on the same day Anthropic laid out a cluster of model mishaps — plus an 8x speed tier for OpenAI's mid-size model and a very big bet on a very young chip startup. The White House is making immediate AI incident disclosure mandatory. The administration's Super Intelligence Force said in a statement shared exclusively with Axios that "SI companies must immediately disclose incidents involving their models and follow with swift, decisive action t