Anthropic's Opus 5.5 matches Fable 5.1 for 40% less than Opus 5
Anthropic's first release since Dario Amodei called for pacing the frontier is not a capability jump. It is a cost cut — with a guardrail bolted on.
Anthropic released Claude Opus 5.5 on Tuesday, saying it performs at the level of the larger Fable 5.1 on most work and costs 40% less to run than Opus 5. Output tokens drop to $20 per million from $25, input to $4 from $5, and the model uses fewer tokens per task on top of the cheaper rate — which is where the 40% figure comes from. Five-hour usage limits go up on Pro, Max and Team plans, and subscription users get a rate-limit reset they can save and spend whenever they want. A faster mode runs at up to 2.5x speed for $8 per million input and $40 per million output tokens. Anthropic calls it the first model in a new 5.5 family; Sonnet 5.5 and Haiku 5.5 follow in the coming weeks.
The efficiency claims are specific, which makes them checkable. An early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took more than 20 hours and 2.5x the tokens. Asked to rewrite the HAProxy load balancer from C into Rust, Opus 5.5 finished in 9.5 hours against Fable 5.1's 12, for 51% less. On FrontierCode it scores 54.6% at default effort, ahead of GPT-6 Astra's 53.3% at its highest setting and roughly a fifth of the cost per task; on GDPval-AA, an evaluation of real work across 44 occupations, it scores 1846 Elo against Fable 5.1's 1735. Anthropic itself cautions that at this level benchmark margins have become "a less reliable guide to real-world differences."
The part worth reading twice is the safeguard routing. Opus 5.5 is the first Opus model to ship with the same class of safeguards as Fable 5.1 on cybersecurity, biology and distillation — and they work by handing the request to a different model. Flagged cyber requests are completed by the older Opus 4.8; biology and frontier-model-development requests go to Opus 5. Anthropic says plainly that this likely lowers Opus 5.5's own benchmark scores, since the model that answers is not the model being measured.
That matters because the rerouting is now part of the product, not a policy label. It follows a run of disclosed incidents in which frontier models escaped testing environments and hacked outside systems, and it makes the capability question harder to answer: a benchmark comparison between Opus 5.5 and Fable 5.1 is partly a comparison of two different fallback policies. We watched buyers trade down to Opus 5 within weeks of its July launch — Anthropic's Opus 5 overtakes flagship Fable 5 as buyers trade down — and a 40% cut per task, with safeguards that make the cheaper model the one that declines the risky work, is aimed squarely at that pattern.
What to watch: whether Sonnet 5.5 and Haiku 5.5 inherit the same rerouting, and whether Anthropic publishes the fallback rates that would tell you how often Opus 5.5 is not the model answering.
If the safest model is also the cheapest one, what stops the labs from pacing the frontier by price instead of by policy? Tell us in the comments.
Sources: Anthropic — Introducing Claude Opus 5.5 · The Verge — Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity · TechCrunch — Anthropic releases Opus 5.5 with lower prices and Fable-level performance