The Take — A refusal rate is a capability score
Artificial Analysis put out a serious benchmark this week and buried its most interesting result in a second chart. I think that is the wrong way round. A model that declines 38 percent of a job does not have less capability on the other 62 — it has a policy. The Cyber Index's headline score turns that policy into a capability gap, silently, and the headline is the thing that spreads.
Here is the mechanism, because this is an arithmetic complaint, not a philosophical one. The index is an equal-weight average of three evaluations: Collinear's 120 audit-and-patch tasks across the OWASP Top 10, Artificial Analysis' run of Vercel's vulnerability hunt scored against an expert-verified golden set, and a 131-task slice of Berkeley's CyberGym where the agent must produce a crashing input and then a patch that survives the project's own tests. No task asks any model to build an exploit — this is defensive work, source-code-in-hand. The honesty is real: the methodology says declines are recorded and reported separately from the score.
Separately is not neutrally. A declined task scores zero in the composite, so the headline arithmetic adds a refusal and a failure together and only one of the two is visible in the number everyone quotes. That is not a small accounting detail here, because the refusal rates are not small: Claude Fable 5.1 declines 38.4 percent of tasks, GPT-6 Astra 37.8, GPT-6 Sol 36.4, Claude Opus 5.5 35.9, Gemini 3.8 Flash 31.8. The two Qwen3.8 models decline more than half — 62.5 percent for the 2.4T model, 56.8 for the 27B. And then the placings: Grok 4.7 and MiMo-V2.6-Pro top the index at 56, GPT-6 Luna is third at 53, and Opus 5.5 sits at 29, with Fable 5.1 and Gemini 3.8 Flash at 25. A 31-point spread between the top row and a Claude model is now the difference between a purchase order and a polite pass.
Our brief on the launch led with the benchmark's own most interesting finding: the frontier models refuse the work. The per-benchmark view is blunter. On the CyberGym third of the index, GPT-6 Sol and GPT-6 Astra declined every task, and Opus 5.5 and Fable 5.1 declined 98 and 99 percent of theirs. Because each evaluation is a third of the composite, GPT-6 Sol's CyberGym third was forfeited before it ran, and the model sits in the thirties on the headline number — an average in which one of three slots is structurally zero. That is the figure procurement will put beside Grok's 56.

The cost column makes the distortion more expensive, not less. GPT-6 Luna scores 53 at roughly 12 cents per task; Grok 4.7 takes the top row at $11.67. So the honest question for a buyer is not how good that 53 is but what it means — and the index cannot say, because the model tying the leaders paused on the exact tasks where its capability is unknown.
Refusal is not inability, and we have the evidence under a different name: Google's Gemini 3.8 Flash Cyber and OpenAI's GPT-5.5-Cyber post scores above 85 on CyberGym's model track, as our brief on Alibaba's 27B security model laid out. The labs can do this work. They sell it under a different SKU, to a different buyer.
What the index measures, then, is a product decision: what a vendor is willing to let a general-purpose API attempt on unvetted prompts. That is legitimate to measure and legitimate to refuse. It is just not capability — and a leaderboard is the format where the distinction disappears, because leaderboards travel as screenshots.
Three reasons that matters more than a chart annotation. First, benchmarks are purchase orders now: enterprise teams use these tables to choose models for security work, and a table that ranks vendors by policy posture without saying so hides its own variable. Second, benchmarks are optimization targets — once a lab tunes against this composite, the only move that reliably gains points is loosening refusals, while the move that gains nothing is tightening them. That asymmetry is built into the arithmetic, not into anyone's bad intent. Third, and this is the part that will get me accused of the wrong argument: I am not asking the labs to unlock anything. I am asking the index to stop reporting a wall as a weakness. If a lab's position is that a vulnerability-patch loop on unvetted prompts is not something a general-purpose endpoint should do, that position is defensible — and the benchmark should record it as a product boundary, in public, in the score itself.
The strongest case for the other side
Give the other side its best version, because it is stronger than "benchmarks are hard."
Refusal here is a deliberate trade made by adults who know their customers, and it is not obviously wrong. The policy that closes exploit-adjacent work on a default endpoint is the one that also keeps a classroom prompt from becoming a working proof of concept, and alignment work is not scoped to what one evaluation asks for — reinforcement learning does not stay in the lane you trained it in. "Reported separately" is the right engineering choice: the score is the score, and the block rate gets its own panel and definition on the same page.
There is also a reading where my complaint is a nit. Vercel's leaderboard says it outright — Anthropic's most capable model is absent because it declines security work, including defensive tasks, and Vercel will add security-enabled versions when the vendors ship them. That is a vendor telling you its number is conditional. And the index claims nothing about total capability: the best model in the suite finds only 41 percent of the expert-verified issues, so the ceiling here is the benchmark's difficulty. On that reading a refusal-shaped zero is indistinguishable from an ability-shaped zero, because everything is low.
Hardest to answer: a refusal is the model working correctly, and it is the one signal here a breach-averse insurer would rather see. Buying cyber agents for a regulated bank, I might put the block rate in the top row rather than a footnote.
Why the take still holds
Because the disclosure exists and the score is still the artifact. I have no quarrel with the block panel; I am saying it should not be the only place the policy shows up. Compose a headline number out of successes and zeros, put it on a leaderboard next to a 56, and that number will be quoted — by sales teams, by procurement, by whoever needs a line for the slide. A metric licensing a policy choice inside a capability ranking is the same failure our AI 101 on what a benchmark actually is warns about: the score is a decision rule dressed as a measurement. If the refusals really are correct behavior, then a ranking that penalizes them penalizes correct behavior — and shown beside a raw capability score, that tension is legible instead of hidden two tabs away.
What would change my mind
- A decomposition. For each of those five models, how many blocked tasks were declined by a provider safety system rather than failing at the harness, and how many were errant blocks on tasks no serious policy covers. If the walls are thin — a few false positives on a narrow class of task — my complaint shrinks to a reporting nit, and I will say so.
- An attempted-only score beside the raw composite. Not a new benchmark, one more column, and it would tell a buyer whether a 25 is a wall or a rope.
- Production evidence. If enterprise deployments of these models, with safety settings relaxed under contract, refuse at the same rate, the wall is real and the index is right to show it. If refusals turn out to be API-default artifacts that vanish for a customer who asks, then the index is ranking vendor defaults and calling it capability.
Until one of those arrives, the fix is cheap and the distortion is expensive: score what a model can do, publish how often it declines, and never let a refusal count as an inability on the same line.
Should a leaderboard rank what a model can do or what it will do? Tell us in the comments.
Sources: Artificial Analysis Cyber Index · Artificial Analysis — Cyber Index methodology · Artificial Analysis — Cyber Index Alliance · Vercel — DeepsecBench · Collinear AI — CWE-bench · Berkeley RDI — CyberGym-E2E