The closed-model premium buys 4.4 months — then it's lock-in

Share
The closed-model premium buys 4.4 months — then it's lock-in

Mozilla's second State of Open Source AI report, out today, puts a price and a shelf life on the thing every enterprise AI budget this quarter is paying for. That is more useful than any leaderboard — and the report is honest about the number it can't measure.

What Mozilla actually measured

The capability gap between the strongest US closed frontier models and the best Chinese open-weights models is 4.4 months, according to the report published September 15. The evidence is specific: Moonshot AI's Kimi K3 scores three points behind Anthropic's Fable 5 on the Artificial Analysis Intelligence Index while costing 30 percent as much. On Terminal-Bench 2.1, run by benchmarking firm Vals AI on a neutral harness so every model faces the same tooling, Z.ai's GLM 5.2 landed within a point of Claude Opus 4.7 and 4.8 at roughly five times less per completed task.

The 4.4-month figure is derived, not measured. Mozilla's CTO Raffi Krikorian grounded it in METR's time-horizon work — the length of task, in expert-hours, a model can finish with a reliable half-success rate — and says the best closed model handles jobs about 1.7 times longer than the best open one. "If the open frontier can handle a seven-hour job, the closed frontier can handle a 12-hour one," he told Ars Technica. "In four months, the open model handles the 12-hour job, and the closed one handles something around 20." Four months is that doubling cadence projected forward. Treat it as the shape of the trade, not a stopwatch reading.

Focused view of a modern data server rack with blinking lights in a blue-lit environment.

The premium is real — inside one narrow band

The practical buying rule falls out of the time horizons cleanly, and it is the most useful thing in the report. No current model, open or closed, reliably finishes tasks above roughly 12 expert-hours. Below eight hours, both kinds work and the cheap one is the sane default. That leaves an eight-to-12-hour band as the entire addressable market for the frontier premium — long-horizon agent work that has to land now. Krikorian's framing: pay when a deadline arrives before the open frontier catches up; skip it for routine work you'll still be running next quarter, "because you'll be able to do it for a fifth of the cost soon, and the model won't be the bottleneck anyway." DoorDash already runs this split, routing routine work to Kimi and saving Fable-class models for the hard cases. Mozilla's earlier survey data showed the same drift: 79 percent of developers use open models, and eight of the ten highest-volume models on OpenRouter by August token count ship open weights.

The cost the benchmark never prices

Model price is the visible number. The expensive one is the harness — the software layer that decides what an agent can see, remember, and do. Closed vendors ship model and harness fused and tuned to each other, so migrating is not a model swap; it is re-architecting the tool your team works inside. Salesforce's Claudeforce, Agentforce and Google's Antigravity are all versions of the same move, and it is the reason a per-token saving can be real and still not worth taking. This is also why the same report that says open models are nearly as capable shows them capturing 4 percent of AI revenue against closed models' 96 percent — a Linux Foundation analysis of May-to-September 2025 traffic that Krikorian concedes is now stale, but whose direction the harness economics explain. The other concentration worth naming: most open models the world actually runs on are Chinese. Krikorian's line is that Chinese labs are running the Android playbook — give the model away, own the ecosystem around it — and he is less worried about the flag than the fact that one country is setting the defaults. His proposed counterweight is the Linux playbook: public compute programs funding fully open reference models, neutral foundations, corporate beneficiaries paying in, philanthropy covering evaluation and audit.

We've watched this economics squeeze from the supply side twice already — China's labs raise API prices while third-party hosts slash them and Z.ai's GLM-5.3-Flash tops benchmarks at one-tenth the price. Now the same logic is arriving inside the enterprise contract.

Where the report gets thin

Mozilla is an advocate, not a referee, and it argues for the outcome it wants. Three things deserve pushback.

The first is that a composite index and a "within one point" result are not the same claim as parity. The jagged-frontier caveat holds: open weights are at or near parity on coding, instruction-following and general knowledge, and the gap concentrates in reasoning, long-context retrieval and agentic work — which is precisely where enterprises want to spend. Mozilla's own inaugural edition, three months ago, reported GPT-5.5 at 83.4 percent on Terminal-Bench 2.1 against 67.9 percent for the best open model then. The headline has moved from a 15.5-point gap to a one-point gap largely because the open side of the comparison changed and the benchmark was run on a neutral harness. Both numbers are true. Neither generalizes.

Second, extrapolation is doing the work in "4.4 months." A doubling cadence that has accelerated before can stall, and the closed labs are not standing still. The same day this report landed, Salesforce shipped Koa with Nvidia — a reasoning model post-trained on the open-weight Nemotron base to handle Agentforce's long-horizon reasoning in-house, per the company's own account of the work. That is the enterprise attack coming from the other direction: open bases, vendor distribution, supplier lock-in intact.

Third, the revenue asymmetry cuts both ways. If open models really were a fifth the cost at near-parity capability for most workloads, the four percent revenue share would be collapsing faster than it is. Something — support, compliance packaging, contractual liability, the fact that nobody gets fired for buying the named vendor — is being bought with the premium, and the report names those things without pricing them.

What to watch

The next edition's gap number is the obvious one, but the more telling signal is whether anyone outside China steps into the open lane at frontier scale — Apertus shows sovereign compute can produce a fully open reference model, not a frontier one. Watch, too, for the first portable standard for agent permissions and the write surface; until that exists, harness lock-in is the frontier labs' most durable moat, and no per-task benchmark will show it.

If a fifth of the cost clears eight of your ten workloads, what's the actual reason your next renewal still goes to a closed lab? Tell us in the comments.

Sources: Mozilla — State of Open Source AI · Ars Technica · Artificial Analysis Intelligence Index · Linux Foundation — hidden economics of open models · TechCrunch — Salesforce and Nvidia's Koa