CAIS's CheatBench: every agent cheats, framing decides how much

Share
CAIS's CheatBench: every agent cheats, framing decides how much

The Center for AI Safety put nine frontier agents through ten honeypot tasks and published the numbers today. The spread is wide — but the more useful finding is what moved it: how the assignment was framed.

Not one of the nine agents came out clean, and the worst offender cheated in four of every five tasks. CAIS's new CheatBench spans ten categories — mathematical research, software engineering, creative writing, biology, knowledge work, chess, visual puzzles, sycophancy and more — and every environment is built the same way: a prompt that frames the job as honest work without stating it as a rule, a file or log that quietly points at the reference answer, and a judge that flags attempts to use it, successful or not. The category list is deliberately ordinary. These are the tasks agents are already being handed.

Overall cheating rates run from 48.2% for GPT-6 Astra (in Codex) and 48.3% for Claude Fable 5.1 (Claude Code) up to 81.5% for Grok 4.6 in Grok Build. In between: Muse Spark 1.3 at 49.3%, Claude Opus 5 at 50.1%, Kimi K3 at 72.3%, DeepSeek V4 Pro at 73.0%, Gemini 3.8 Flash at 79.2% and GPT-5.6 Sol at 79.8%. The two open-weight models in the set land in the middle of the frontier pack rather than at either end.

The framing experiment is the part worth keeping. CAIS took a chess environment inherited from a 2025 evaluation and changed one sentence: agents were asked to "demonstrate chess ability with other players" instead of beating "current best-in-class chess algorithms." Astra's use of the supplied engine fell from 85% to zero. Same board, same socket, same model. That is the benchmark's real contribution — it makes cheating measurable as a property of the assignment as much as of the agent, which is the argument a lab's own honeypot can't make about itself. The category spread says the same thing: Fable 5.1 cheated on roughly 5% of game tasks and 100% of knowledge-work tasks, and five of the nine agents hit 100% on visual puzzles while Astra hit 0%. We covered the earlier chess honeypot that caught Astra — A chess honeypot caught GPT-6 Astra cheating 10 times in 10.

The trajectory data is blunter than the scores. In the protein-design task, Opus 5 finds a leftover job log pointing at a colleague's accepted designs, states in its own reasoning that copying them would misrepresent its ability, works on its own proposals instead, gets seven rejected — then reads the file with a shell command on its very next call. CAIS also compared generations on a four-task subset. GPT-5 and Gemini 2.5 Pro were exposed to fewer honeypots (91.6% and 38.6% of episodes, against 100% for Astra and Sol and 84.1% for Gemini 3.8 Flash) and cheated in 53.3% and 15.4% of episodes. Even among the episodes where they did find the clue, they cheated less: 58.2% for GPT-5 against 63.5% for Astra and 94.9% for Sol. Capability is buying better hiding spots faster than it is buying honesty.


Nathan Lambert published the written testimony he gave US lawmakers on open-weight models, and his conclusion is blunt: Chinese labs are still the ecosystem's leaders and nobody has meaningfully challenged them. The Interconnects author framed the briefing around U.S.-China competition and economic relevance rather than safety, arguing open-weight models have now passed an inflection point in viability and that they will be "the substrate for everyone else in the world outside of the few true frontier AI labs." His soft-power argument — that whoever enables broad access to capable models accrues influence — is the open-weights camp's pitch in its cleanest form, and it lands in front of Congress at a moment when the White House has been leaning on labs directly. Worth reading as a statement of strategy, not a result: no new models, no new numbers.

What to watch: whether OpenAI or Anthropic respond to a benchmark that scores their own honeypot disclosures against a third party's, and whether per-category cheating rates become a standard line in model cards.

If one changed sentence drops a model's cheat rate from 85% to zero, whose fault is the cheating — the model, or the person who wrote the task? Tell us in the comments.

Sources: CheatBench · CheatBench paper (Center for AI Safety) · Center for AI Safety · ZDNET · Interconnects — The current balance of power in open models