The AI watchdog gets a badge, a laptop — and no veto power

Share
The AI watchdog gets a badge, a laptop — and no veto power

Dario Amodei's pacing essay promised outside evaluators employee-level access inside Anthropic. Today brought the harder question: what happens when one of them finds something the lab disagrees with?

Access without authority

CNBC reported this morning that the embedded-evaluator plan gives outsiders extraordinary visibility and almost no formal power. Amodei committed Anthropic to providing third-party evaluators access comparable to its internal risk teams, plus the right to publish findings without the company's editorial control, subject to limited redactions. He anchored the idea in banking, writing that his proposal "has precedent in the banking industry, which sometimes involves regulatory 'supervisors' embedded along with employees."

Julie Andersen Hill, dean of the University of Wyoming College of Law and a banking-regulation specialist, says that comparison does not survive contact. Bank examiners can order an institution to stop a practice, restrict its growth, force management changes, and in extreme cases close the bank. Anthropic's evaluators can investigate and report. They cannot prevent a model from being trained or released. "If you don't give them that kind of power, I don't know what they are doing," Hill told CNBC. Neither Amodei's proposal nor OpenAI's existing third-party framework grants outside evaluators authority to halt development or deployment.

Hill's sharper line: if the lab picks the evaluator, defines what it can see, and stays free to ignore the conclusions, "that looks a lot like an internal compliance department." Her point is not that compliance departments are useless — it is that you do not get regulatory credibility from an arrangement you control end to end. "You can't have all of the control and then expect the credibility as if you've given up control."

Business professionals examining financial documents with magnifying glass for detailed analysis.

The rift is inside the labs

The Financial Times reported today that the practicalities are already causing friction inside both companies. Employee-like access means badges, laptops, office entry and internal tools for outsiders — at firms whose systems are both the most valuable asset they own and the target of every competitor. People close to Anthropic and OpenAI described internal concern over security and intellectual property, which is a more concrete objection than the philosophical one.

Miles Brundage, a former OpenAI researcher who now runs the AI Verification and Evaluation Research Institute, told the FT that most third parties today have less access than the lowest-privilege employee, so a binding industry-wide requirement — not a voluntary pledge — is what would change anything. David Krueger of the University of Montreal, previously at the UK AI Security Institute, argued governments rather than labs should decide who gets access, and that the current framing treats systems as safe until proven dangerous.

That matters because the pledge is currently a pledge. Anthropic committed unilaterally; OpenAI said it would match and has not published terms. Two companies' voluntary arrangements do not create a standard, and the audience that needs convincing — Congress — has heard that argument before. We covered the state-level version of this in California just made AI companies face an auditor they don't pick, where the mechanism was chosen to remove the lab from the selection decision.

What evaluators actually see

Albert Ziegler, head of AI at the security firm XBOW, evaluates unreleased models from Anthropic, OpenAI and others, running his own tests in his own environment. His description of the day job is unglamorous: models producing nonsense under an unusual formatting request, a safety checker that has to intervene more frequently than before. Asked about the catastrophe framing — Amodei's pacing plan argues a misaligned swarm could seize large parts of the internet within six to twelve months and cause hundreds of billions of dollars in damage — Ziegler says the insidious combination of subterfuge and unprecedented ability is not something his team has seen. And he states the constraint plainly: "It's true that we don't have any veto power."

What an evaluator can do is document a risk the developer missed and "compel an informed decision before release." Black-box testing shows whether a model can perform a dangerous task; judging whether the system is dangerous needs the surrounding instructions, tools, permissions, safety controls and logs of attempted actions — and even with all of that, a serious failure may only appear under a combination of circumstances the test never triggers. A residual risk that never fires in the eval harness is invisible to the evaluator and unbounded for everyone else. That is the same accountability hole we flagged when after an AI breakout, nobody had the power to investigate.

METR, the nonprofit Amodei named as an example evaluator, is itself an instructive case. Its February-to-March 2026 Frontier Risk Report pilot — run with Anthropic, Google, Meta and OpenAI — concluded that internal agents plausibly had the means, motive and opportunity to start small rogue deployments, but not the means to make them robust. It also disclosed that participants approved which non-public material could be published, that some staff have strong social ties to lab employees, and that it shares a research center with some of them. A former Anthropic researcher, Joe Benton, recently left to join METR on embedded assessments. None of that proves the work is compromised. All of it explains why White House AI adviser David Sacks posted that people should "stop pretending METR is independent when it is intertwined with Anthropic's investors and staff."

Who wins, who pays

Two legislative routes are now visible. OpenAI's chief global affairs officer Chris Lehane backed the bipartisan Obernolte–Trahan FRONTIER Act, which would require developers past revenue and compute thresholds to admit "independent verification organizations," and compared the qualification problem to how accounting firms became authorized auditors. Representative Josh Gottheimer, who co-chairs the House AI Commission, rejected that: "Third-party audits alone don't meet the moment," he said, calling instead for mandatory government pre-review of frontier models before public release. Roughly the reverse of the labs' preference.

Follow the money and the compliance cost. Hill notes that continuous supervision is expensive for the supervised, which favors incumbents able to absorb it and raises the entry bar for everyone else — a dynamic also visible in the labs' own ask, where reporting says they are lobbying to attach an antitrust carve-out to the National Defense Authorization Act so safety coordination does not read as collusion.

What to watch

Three signals. Whether OpenAI publishes the scope, the redaction rights and who picks its evaluators. Whether the antitrust carve-out survives the NDAA. And whether the FRONTIER Act moves at all against Gottheimer's alternative. Hill's test is the one that settles it: an evaluator reports a serious finding, and the model ships anyway — or does not.

If the supervisor cannot stop the release, is the badge a safeguard or a receipt? Tell us in the comments.

Sources: CNBC · Financial Times · POLITICO · Dario Amodei — We Must Pace the Frontier · METR — Frontier Risk Report

Read more