Google's fix for agents that memorize their tests

Share
Google's fix for agents that memorize their tests

Let an agent rewrite the scaffolding around its own model against a fixed set of test tasks and it learns the tests, not the job. Scores on the tasks it practiced climb; gains on anything new shrink or vanish. Google Cloud AI Research's answer, written up by The Decoder today, is not to lock the scaffolding down but to regulate how the search moves through it — and the numbers suggest the regulation, not the model, is what makes self-improvement stick.

The catch in recursive self-improvement

The scaffolding is the harness: the prompts, control flow, tools, memory and context management wrapped around a frozen model. Which file the agent reads before editing it, how it recovers from a mistake, what it does with its results — that is harness design. A growing share of recent agent progress comes from there rather than from new model weights, and until recently humans did it by hand: watch a failed run, patch the workflow, repeat.

Newer methods automate the loop — a language model proposes harness edits, the system runs them against a benchmark, and the winners stay. The RRSI authors call this a practical form of recursive self-improvement: the system produces feedback it then uses to change its own behavior. Their point is that the loop has a failure mode baked in. Because the same limited set of tasks grades every candidate, the search memorizes — it finds patterns that fit one benchmark, favors edits that scored well by luck, and piles on complexity that raises the test number without making the agent better. The paper, submitted to arXiv on September 21 and covered by The Decoder on October 4, names the fix: Regularized Recursive Self-Improvement of Agent Harnesses, or RRSI.

A leash on both ends of the loop

RRSI leaves the harness fully editable and constrains the search instead. On the proposal side, a budget caps how many independent edits one candidate may bundle, and the budget shrinks over time — big rewrites early, small traceable changes later. The proposer is fed the full history of past attempts so a hypothesis already falsified does not get redrawn, and when progress stalls it is pointed at harness components no one has touched yet. On the selection side, a critic screens every candidate for suite-specific tricks — hardcoded task names, benchmark-shaped shortcuts — before it is ever evaluated; a floor rejects any gain smaller than the evaluation's own noise; a cost rule only accepts higher token spend when a measured improvement pays for it; and components that stop helping get pruned.

The results are the argument. With Claude Opus 4.8 frozen across eight benchmarks — coding (Terminal-Bench 2.1), document work (Harvey LAB) and engineering design — RRSI gained up to 14.1 points on the split it evolved against and up to 4.7 points on five benchmarks it had never seen, the biggest held-out gain landing on JobBench. It ran on about 30 percent fewer policy tokens than unregularized evolution. Of the methods the authors compared against, RRSI posted the smallest training gain and was the only one that stayed above the plain baseline on every unseen benchmark; two rivals ended up below it. There is also a transfer result worth pausing on: a coding harness tuned with Gemini 3.5 Flash lifted the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points with no modifications. The mechanisms that fell out of the search do not depend on the model that found them. The code is public on GitHub, with candidate harnesses evaluated in git worktrees and an edit history that records hypothesis, score and verdict for every accepted change.

The benchmark is becoming a negotiation

This matters because the harness is now where leaderboard numbers are made. Today we covered Kaiming He's harness giving Claude a perfect ARC-AGI-3 score, and in August NVIDIA's AVO agent cleared every ARC-AGI-3 public level by moving Claude Opus 5 from a roughly 30 percent model baseline to 100 — again, harness, not weights. The Decoder reports the same fragility from the other direction: a purpose-built harness took Claude Opus 4.6 to 97.1 percent in a familiar environment and 0 percent in an unfamiliar one. Once harness search is automated against public benchmarks, any reported score is a joint product of the model and the search budget behind it — and only one of those two is usually disclosed.

Who wins: teams with held-out evaluation discipline and production bills to cut — the 30 percent token reduction is money, not a paper metric — and Google Cloud Research, which gets to publish the guardrail for a technique it also benefits from. Who loses: anyone reporting in-distribution gains as capability. This is also the concrete, adoptable counterpart to the worry in Deep Dive — The intelligence explosion, in the labs' own numbers: the recursive self-improvement those 22 authors warn about is already running at the system level, and here it comes with a config file instead of a manifesto.

What skeptics say

The strongest pushback is not from Google. An independent evaluation, "Rethinking the Evaluation of Harness Evolution for Agents," concludes that under matched feedback and inference budgets on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, automatic harness evolution does not consistently beat simple task-level search and shows limited generalization — its authors argue most published harness-evolution results share the same benchmark between search and final scoring, so the gains can't be attributed to better harnesses at all. RRSI answers the protocol critique with held-out splits, but it carries its own soft spots: the study covers frozen models only, so it says nothing about weights that change underneath; the critic policing "benchmark-specific tricks" is itself an LLM, and anything with a denylist can be gamed around; the best held-out gain is 4.7 points; and the unmodified baseline harness still uses fewer tokens than any evolved variant. The honest reading is that RRSI shrinks the gap between practiced and unseen performance — not that self-improvement is solved.

What to watch

Whether held-out evaluation becomes the reported standard for harness-evolution work — the protocol fight decides how much of this literature is real. Whether the Apache-licensed repo gets adopted outside the paper. And the siblings: Nvidia's SoL-Pi cut coding-agent token use by up to 49 percent by rebuilding the harness, Google DeepMind's Dream RSI improves search agents by replaying past attempts. When optimized harnesses start shipping as products, "which model is best" becomes the wrong question — the answer will ship with a benchmark history attached.

Sources: RRSI code (GitHub)

Read more

Jay Clayton will run Trump's Super Intelligence Force

Jay Clayton will run Trump's Super Intelligence Force

Washington gave its AI rebrand a name and a clock this Sunday, SoftBank's founder asked the world to take superintelligence off autopilot, and Google published a fix for agents that quietly memorize the tests they are graded on. Donald Trump named director of national intelligence Jay Clayton as the White House's AI czar on Sunday, putting him in charge of a "Super Intelligence Force" tasked with reporting on AI risks and opportunities within 120 days. Clayton keeps his day job running the inte

Bessent compares AI chiefs calling for rules to Hannibal Lecter

Bessent compares AI chiefs calling for rules to Hannibal Lecter

Washington is done listening to the labs warn about themselves — the administration now answers doom headlines with a shrug, and Saturday brought the sharpest version yet from the Treasury Secretary. Treasury Secretary Scott Bessent says the AI executives asking government to step in should just slow down instead. On The Axios Show Saturday, co-founder Mike Allen asked whether the doomsday fears now coming from lab leaders amount to "BS." Bessent allowed that "we have to be prepared for every o

How OpenAI's own models helped build its Jalapeño inference chip

How OpenAI's own models helped build its Jalapeño inference chip

OpenAI is handing chip design to its own models, Google is pausing an open-source security program under a flood of AI-written reports, and a new benchmark puts numbers on how Chinese models handle political taboos. OpenAI's hardware team used the company's own models to help design Jalapeño, the custom inference chip it co-designed with Broadcom — and the models cut real work, not just slideware. In a Q&A published Sunday, VP of Hardware Richard Ho told Ian Cutress that in one example the inte

Open Source Radar — October 4: Cloudflare's OS for agents

Open Source Radar — October 4: Cloudflare's OS for agents

Today's trending signal is the toolchain opening up: Cloudflare handed the public its internal AI workspace, Addy Osmani's engineering skill pack crossed 100,000 stars, and a nonogram benchmark is puncturing model confidence in public. Cloudflare OS (TypeScript, ~10,700 stars, Apache-2.0) — Cloudflare just open-sourced the "AI productivity environment" a large share of its own workforce uses daily, and it reads less like a demo than an internal product with a security team's fingerprints on it.