Over 100 AI experts attach conditions to the labs' evaluator pledge

Share
Over 100 AI experts attach conditions to the labs' evaluator pledge

The people who would do the auditing just told the labs what independence has to mean — and one of their conditions is protection from being sued for it.

More than 100 AI researchers and evaluators published a letter Friday setting minimum conditions for the third-party evaluation that Anthropic and OpenAI have promised, with signatories including Geoffrey Hinton, researchers from Johns Hopkins and Stanford, and staff at the nonprofit evaluator METR. The letter was organized by the AI Evaluator Forum consortium and shared first with CNBC, and it does not thank the labs so much as invoice them: to count as credible, embedded evaluation needs "scientific objectivity, transparency, independence, and robust protections against interference," spelled out in five conditions.

The conditions are specific enough to be uncomfortable. Evaluators must not be owned or governed by a frontier lab, must not hold other significant commercial business with one, and must not take any payment contingent on their findings. Labs are to embed multiple evaluators across risk areas rather than one house reviewer. NDAs must be narrowed so evaluators can talk to the boards of the companies they audit and release findings publicly, limited only by a time-boxed redaction process for IP, customer data, privacy, security, and safety. And labs must shield evaluators "from retaliation," including retaliatory litigation, with funding arrangements that hold up even when the findings are unflattering.

That last one is the clause that converts a press release into a contract. "We're just trying to really demonstrate a shared common ground on basic principles and ensure that independent oversight can be a meaningful tool for managing AI risk broadly," Conrad Stosz, chair of the AI Evaluator Forum, told CNBC. The letter asks for access equal to that of the labs' own most privileged employees — the same systems, data, tools, physical spaces, and unscripted one-on-one time with staff. Vinh Nguyen, a Council on Foreign Relations senior fellow and former chief AI officer at the NSA, framed the stake plainly: when a few labs control capabilities that can endanger critical infrastructure, "the government and the public cannot be dependent on those labs' own account of what's secure and safe."

This is the second half of a story that moved fast last week. Dario Amodei proposed employee-like access for outside evaluators, and Sam Altman agreed within hours — the pledge we called the industry's new consensus, and one that still had no named evaluators, no defined scope, and no timeline when Altman matched Amodei's evaluator pledge. The specifics of what Amodei was actually offering are in Amodei's pacing plan puts outside auditors inside Anthropic. TechCrunch spent this week asking both companies which evaluators, how many, with what access, disclosable to whom — and got no answers. The letter is the outside world writing the terms the labs left blank.


Disney hired Karandeep Anand, until recently the chief executive of Character.AI, as its first chief technology officer — a newly created role reporting directly to CEO Josh D'Amaro, effective October 2. Anand will oversee enterprise technology, infrastructure, data and AI platforms, plus product and engineering, and Disney is hiring members of Character.AI's technical team along with him. The company's stated priority is "technology in service of creativity," with Disney+ positioned as an integrated ecosystem of streaming, parks, merchandise, and games.

The hire is a little awkward on paper. Disney has clashed with Character.AI before — it sent the startup a cease-and-desist letter over unlicensed use of its characters, which Character.AI says it complied with. Now Disney is bringing over the executive who ran that platform and staffing up around him. A media giant's first CTO is not an infrastructure purchase; it is a bet that the consumer-AI playbook — engagement, character IP, interaction at scale — is the next thing Disney sells.

If embedded evaluators can't be paid in a way that survives a bad finding, is the access worth anything? Tell us in the comments.

Sources: AI Evaluator Forum — Minimum Conditions for Embedding Evaluators · CNBC — Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter · Quartz — Over 100 AI experts are warning that safety evaluators aren't truly independent · TechCrunch — Anthropic and OpenAI want to embed safety evaluators. Will they really be independent? · CNBC — Disney names CTO for the first time as media giant expands tech push · Variety — Disney hires former Character.AI CEO Karandeep Anand as its first CTO