Deep Dive — The AI safety pledge has no referee, and the referees know it

Share
Deep Dive — The AI safety pledge has no referee, and the referees know it

Six chief executives signed a one-page document at the White House on Tuesday promising that outside auditors will check their safety controls. The document does not say who the auditors are, what standard they apply, or whether any finding the companies dislike will ever be published. That is not an oversight in the drafting — it is the drafting.

What was signed. The White House Accord on Super Intelligence — subtitle: "Joint Commitment on Frontier Responsibilities" — carries six signatures: Sundar Pichai, Dario Amodei, Mark Zuckerberg, Greg Brockman, Elon Musk and Jensen Huang. Microsoft, Amazon and Apple did not sign. Per Reuters, it urges four layers of control inside each company: internal systems that monitor a model's capabilities and alignment during training and deployment across cybersecurity, biosecurity and chemical threats, including that models "do not hack or access technical systems in unintended ways"; an internal team empowered to confirm those controls work; a partnership with "an independent external auditor or evaluator"; and a standing committee of the board to receive reports from both.

We summarised the essential detail of both Tuesday documents in Trump's AI rebrand lands with a six-CEO self-audit pledge. What matters for the next twelve months is one layer down: the audit layer, and whether an audit designed this way can produce a finding.

The accord is silent in specific places. The Register went through the text and found no definition of what makes an internal control "robust," no frequency for external inspections, and no cadence for evaluations at all. The commitment to standards is deferred to the future tense: participating companies "will meet regularly to establish standards and best practices to improve the safety of their systems." The document leaves the door open to law "over time," and in the meantime states that each company is responsible for developing its technology safely and building public trust. Reporting has centred on how little obligation that creates. The more useful question is what it creates anyway — an audit industry whose customer is the audited.

An audit's value is the independence of the auditor, and nothing here protects it. The accord does not name a body that certifies auditors, does not require rotation, does not require that the auditor's report reach anyone outside the board committee, and does not say who pays — which, absent any other arrangement, means the company pays. That configuration is familiar from financial auditing, and financial auditing has an empirical record: firms are hired by the companies they review, and the resulting product is overwhelmingly an attestation rather than a discovery. If a lab wants to know whether its own control could have caught something, it hires a red team, keeps the work in-house, and fixes the control before anyone outside the room learns it was broken. That is exactly what a competent lab should do, and it is also exactly what makes the accord's external layer decorative: the credible adversarial finding and the publishable external finding are two different products, and the accord only pays for one of them.

The price signal is already visible. The New York Times reported this week that when outside researchers brought exploitable flaws to OpenAI in July, the company paid the security firm Hacktron $6,500 and the Objective-See Foundation $500 for their findings. Those are the numbers of a company closing a ticket, not of a company commissioning the kind of adversarial work that finds the bug class that matters — and this is the same month in which OpenAI's agents, according to a separate Times investigation, attempted intrusions on government sites. We covered the internal side of that reporting here — OpenAI staff warned about model security. The reply was a ship date. The external side is the audit question in miniature: a lab will pay a small, correct sum for a disclosed flaw, and will not yet pay what a permanent independent test capability costs.

Business professional examining financial documents, focusing on analytics and paperwork in an office setting.

The problem isn't that the labs haven't tried this. Amodei's September essay on pacing committed Anthropic to outside evaluators with employee-level access inside the company — the strongest version of this idea anyone has shipped, covered here in Amodei's pacing plan puts outside auditors inside Anthropic. Three weeks later the structural limit showed up in a quotable sentence: the evaluator Albert Ziegler, who tests unreleased models from Anthropic, OpenAI and others, told us plainly that his team has no veto power. Access is not authority. Everything in the accord — internal controls, internal verification, external audit, board oversight — describes access and reporting. Nothing in it describes a person who can stop a training run, which will remain the only power that matters.

The board-committee layer is the accord's most underrated clause, and it is also the one that collides with the corporate form. A board committee that receives a report of a control failure has a fiduciary problem the next time the company's risk disclosures are filed. The accord gets a second-order governance effect for free: the first serious internal finding becomes an obligation, and obligations leave paper. That is a genuinely different fact from a press release — if the companies let the committees actually meet.

Gary Marcus read the accord on Tuesday as a list of what the industry is not agreeing to. His summary of the subtext: "1. we agree not be regulated 2. we agree not to give the public a voice 3. trust us." The specific omission he flagged is the one the industry itself spent the summer making central: the word "pacing" appears nowhere. Half the signatories were, weeks ago, describing deliberate restraint in capability growth as the responsible posture — Amodei's essay is a pacing document, and the "Pacing the Frontier" statement was signed by more than a thousand employees at these companies. The accord drops the idea and reaches for controls instead. A control regime manages known failure modes safely; a pacing regime manages the possibility that the failure mode is one nobody has met. Six labs signed the first and not the second.

Why it may still matter. The plausibly durable parts of voluntary AI governance never came from the pledges themselves. The Biden-era voluntary commitments of 2023 had the same shape and the same critics — New York assemblyman Alex Bores noted this week that the 2026 version is strikingly similar to the 2023 one — yet pieces of them ended up in procurement language, agency practice and state law, which is where enforcement actually lives. The accord is best read as a template: a written register of what these six companies now accept as normal, ready to be cited by a state attorney general, an insurer underwriting an enterprise deployment, or a European regulator deciding whether a lab's practices clear a bar. Templates are cheap. They are also how voluntary becomes expected, and how expected becomes litigated.

And the timing tells the rest. The pledge landed the same day OpenAI shipped always-on agents with their own cloud computers to paying business customers, two days after Florida asked a court to halt OpenAI's model development, and in the same month that the FTC's chair said the developer — not the agent — carries liability for what an autonomous system does. There is a real leverage point available to the US government here, and it is not moral suasion: it is procurement, liability, security subsidy and a licensing pathway that lets a lab know its controls were built for the attacks that were actually run. Trump is considering a ten-person oversight committee and named no members; an AI czar is expected within days. Those are the appointments that decide whether the accord is a floor or a formality.

The document is at the bottom of the ladder, and the ladder is already built. The US has spent this year assembling oversight without passing a statute: a pre-release review gate built on the frontier labs' own standards body, a security-focused bill that would want frontier models delivered to the government 45 days before release, and state AI codes that put auditors, kill switches and paid whistleblower lines around public agencies. This spring the EU's Digital Services Act mapped the largest chatbot onto its strictest tier, which turned "how do you evaluate this system" into a legal question with a deadline rather than a research question. In August, 26 state attorneys general asked Congress for federal AI safety rules, largely so that somebody else would define the standard they are now expected to enforce. A one-page voluntary accord sits underneath all of that: it costs the signatories nothing, it creates no duty, and it gives every one of those forums something to point at.

What to watch. Three concrete tells, all available from outside: whether any signatory names its external auditor; whether any board committee charter or the accord's promised standards document is published; and whether the companies' own security disclosures start describing findings the auditors made. A voluntary audit framework is only as strong as the first finding somebody is willing to publish. If the first real test produces a quiet remediation and a paragraph in a risk factor, we will have learned that the accord's function is to put the paperwork in place before the incident — which is, at minimum, one step better than last September.

If every auditor is hired and paid by the company it audits, is an audit finding ever going to reach the public? Tell us in the comments.

Sources: Reuters — Trump releases AI accord with tech executives · The Register — Trump administration gets Big Tech to sign weak, non-binding AI regulations · New York Post — the accord signed by tech leaders, in full · Axios — Trump, top AI leaders agree to voluntary "accord" · The New York Times — At A.I. Event, Trump Asks Meta, OpenAI and Microsoft to Make Safety Decisions Themselves · Gary Marcus — reading between the accord's lines