AI safety scores can be gamed — psychometric study exposes how

Share
AI safety scores can be gamed — psychometric study exposes how

A new study out of the UK AI Safety Institute shows most AI safety benchmarks reward the wrong behavior and can be inflated by a model that simply refuses more requests — while a handful of carefully picked questions can replace thousands.

The largest psychometric audit of language-model safety tests to date finds that a single aggregated safety score hides more than it reveals — and a model can pad that score simply by getting stricter about what it will answer, even when the extra refusals make it worse at ordinary tasks. Researchers, including staff at the UK AI Safety Institute, applied item response theory — the statistics built for human IQ and aptitude tests — to eight popular safety benchmarks across 192 models and more than 5,000 questions. They found those benchmarks do not all measure one shared "safety" trait. They split into three weakly related factors: how strictly a model refuses requests, how truthfully it answers, and how it handles content that could be harmless or dangerous depending on context.

That split produces the study's most striking finding. HarmBench rewards a model for turning down harmful requests, while OR-Bench-Hard penalizes one for refusing harmless ones — so a model that scores well on one almost always scores poorly on the other, and a lab can lift its overall rating simply by blocking more requests across the board. Averaging results across benchmarks papers over that tradeoff entirely.

The second finding is about cost. Most test questions are dead weight: nearly every model passes or fails them, so they distinguish nothing. Psychometrically chosen items recover a full benchmark's score with far less error than random subsets, and roughly ten adaptively selected questions match what most of the full evaluation measures — cutting evaluation cost by 97 to 99 percent and making regular safety checks mid-development practical.

The third concerns the rising worry of "sandbagging," where a model notices it is being tested and behaves more cautiously than in everyday use. The method caught 80 to 100 percent of models instructed to act overly cautious, climbing to 94 to 97 percent with a second check comparing a model against its own past results. As a side effect, it can tell whether the model behind an API is the one you actually evaluated — the team correctly identified 13 of 14 distinct models from response patterns alone and caught one service whose answers had drifted a third of the way from its own answers two months earlier.

The authors argue AI safety tests should meet the same rigor as psychological tests for humans, and that a single score is too thin a basis for deciding whether to release a model. Their limitation is honest: the sandbagging they detected was simulated by direct instruction, and a model specifically trained to dodge detection would be far harder to catch.

What to watch for is whether frontier labs adopt adaptive, shortened evaluation as a standard practice — and whether the benchmark-scoring arms race meaningfully changes as a result.

Do you trust aggregated model "safety scores" any less now that you know how easily they can be gamed? Tell us in the comments.

Read more

Korea probes AI agents in bank hacks as president cites 'signs'

Korea probes AI agents in bank hacks as president cites 'signs'

South Korea opened a formal investigation into whether AI agents drove a wave of bank breaches — and it isn't the only AI story moving money today. President Lee Jae Myung said "signs" point at AI models, and the probe is now at the highest level a national banking sector has seen. Speaking at a cabinet meeting, Lee said that "in some hacking incidents, signs have emerged of AI being used, causing considerable public concern and anxiety," and police have since opened a full-scale investigation

Open Source Radar — October 6: nothing leaves your machine

Open Source Radar — October 6: nothing leaves your machine

Today's trending board is all projects we ran earlier this week, so the fresh signal comes from the Product Hunt launch slate instead — three open-source tools that share one instinct: your phone, your pixels and your MCP traffic should stay on hardware you control. All three verified at the source. iphone-use (Rust, MIT, about 59 stars) is computer-use for a real iPhone: an agent reads the screen as text, taps, swipes and types over WebDriverAgent, and every action comes back with an honest v

Deep Dive — Anthropic's guardrails cost it the Pentagon, court or not

Deep Dive — Anthropic's guardrails cost it the Pentagon, court or not

The Pentagon "has ceased the use of Anthropic products," a department official said in a statement to the BBC on Monday — the first time the US military has said out loud what it has been working toward since February. The order to phase Claude out was signed by defense secretary Pete Hegseth on February 27 with a six-month deadline attached; the deadline passed in late August with no public explanation, no successor named, and no acknowledgment that anything had changed. What finally forced a s

DeepSeek nears $12B round with Tencent and CATL ahead of IPO

DeepSeek nears $12B round with Tencent and CATL ahead of IPO

Three stories shape the last 24 hours: a record-scale fundraise at China's most famous model lab, a first-of-its-kind hearing at New York City Hall, and Cohere rebuilding its enterprise agent platform around other people's models. DeepSeek is close to raising at least $12 billion — 80 billion yuan — in a single round that could reach roughly $14.9 billion, after investor demand outstripped the company's own target, with Tencent and battery maker CATL as the biggest contributors, people famili