OpenAI halves GPT-6 Sol and Luna prices hours after Anthropic's cut
Two frontier labs spent Tuesday afternoon arguing about the same thing — not who is smartest, but who costs least per task.
OpenAI cut API prices for GPT-6 Sol and Luna in half on Tuesday, hours after Anthropic cut Opus 5.5 pricing by 40%. Sol drops to $2 per million input tokens and $10 per million output, from $4 and $20; Luna goes to $0.10 and $0.50 from $0.20 and $1.20, measured against the promotional pricing of the GPT-5.6 versions they replace. OpenAI attributes the cut to better prompt caching and cheaper inference rather than to smaller models — Sol and Luna are trained with the same methods as Astra, the flagship it shipped on September 3, and cached input reads now carry a 90% discount. The models are live in ChatGPT Work and Codex for paid users, with Free and Go users getting Luna in the desktop app; they are not yet in the main ChatGPT app.
The reliability claim is the more interesting half of the announcement. On OpenAI's internal factuality evaluation — built from de-identified ChatGPT conversations where users had flagged a factual error by an earlier model — GPT-6 Sol makes about half as many mistakes as its predecessor, which OpenAI describes as approaching Astra-level reliability at much lower cost. Luna, at higher effort settings, matches GPT-5.6 Sol at roughly a hundredth of the price. These are vendor-run evaluations on a deliberately hard slice of traffic, so read the ratio as a direction, not a spec.
OpenAI's competitive numbers are aimed squarely at Anthropic. On AutomationBench, a test of business workflows across 47 tools, OpenAI says GPT-6 Sol at extra-high effort scores 33.2% at $0.27 per task, against Claude Opus 5 at maximum effort scoring 26.9% at 11.1 times the cost. On DeepSWE, a long-horizon software engineering benchmark, Sol at maximum effort lands 68.8% — within 1.1 points of Fable 5's 69.9% at roughly 80% lower cost per task. The competitor figures come from published reports rather than OpenAI's own runs, and the footnotes do real work here: OpenAI notes that Fable 5.1's reported cost omits the Opus 5 fallbacks that fired on about 40% of tasks.
The timing is the story. We covered Anthropic's cut when it landed Tuesday afternoon — Anthropic's Opus 5.5 matches Fable 5.1 for 40% less than Opus 5 — and both releases arrived wrapped in safety framing about pacing the frontier. Within hours, both labs had moved on price instead, and the axis of competition is now cost per task on evals neither lab wrote.
OpenAI also published the rules it wants outside safety assessors to work under. The company named four priority areas for independent assessment — safety cases across training, evaluation and deployment; the safeguards themselves, including misalignment monitors; the capability evaluations behind its Preparedness Framework; and independent investigation of misalignment incidents — alongside principles for how assessors should operate: claims pre-registered and scoped before work begins, proportionate access to internal systems, disclosed conflicts of interest, enforceable confidentiality, and time for the lab to fix findings before publication. Bloomberg reported the plan the same day. The gap is enforcement: nothing in the document obliges OpenAI to accept a finding it disagrees with, or to publish one.
What to watch: whether Anthropic answers with another price cut, and whether either lab starts publishing cost-per-task figures on evaluations it did not run itself.
If frontier work is now priced like a commodity, what happens to the labs' safety arguments? Tell us in the comments.
Sources: OpenAI — Introducing GPT-6 Sol and Luna · TechCrunch — OpenAI launches GPT-6 Sol and Luna · OpenAI — Priorities and principles for effective third party assessments · Bloomberg via Techmeme — OpenAI plans third-party safety evaluations