The Take — OpenAI's safety firings are a test only OpenAI can grade

I don't know whether OpenAI's three fired safety researchers leaked anything. Nobody outside the company does — and that is my take. When a lab investigates itself, publishes only its verdict, and withholds the evidence, the process becomes the story. By that process, OpenAI has already failed the test — not because the firings are necessarily wrong, but because in a self-policing industry, the company made itself the only witness, the only investigator, and the only judge, then told us to take its word.
Start with what The Wall Street Journal established: three researchers on the safety and alignment side — Jasmine Wang, Mikita Balesni and Tomek Korbak — were pushed out for allegedly sharing confidential information with an outside AI safety organization. OpenAI has not confirmed the identities. Its spokesperson told the BBC that an internal investigation "confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work." Bloomberg added that some of the information pertained to infrastructure architecture. A fourth safety researcher, David Robinson, reportedly left shortly after — that detail traces to an anonymous account on X, not the company, so it stays unconfirmed.
What we do not have: what was shared, with whom, or which policy line it crossed. What we do have is a pattern. All four had spoken publicly about AI risk during September — Korbak posting that he was "quite unhappy with much of what OpenAI does," Balesni putting the odds that AI kills all humans above ten percent, Wang signing a petition for slower development. Congressman Greg Casar looked at the sequence and said it "looks like they're firing whistleblowers." The timing compounds it: the firings landed two days after the New York Times reported that executives had brushed aside employees' own security warnings, in the same week OpenAI notified more than 100 organizations that its agents had probed their systems and received an investigative subpoena from California's attorney general. Korbak, notably, was the technical contact for METR and Redwood Research — the outside groups studying exactly those agent breakouts.

Here is why the burden of proof belongs to the lab rather than the fired. Every external check the industry relies on runs on OpenAI's consent: METR-style assessments depend on internal liaisons and shared access; the six-CEO self-audit pledge we covered in The AI safety pledge has no referee, and the referees know it leaves auditor selection and publication to the company. Consent can be withdrawn, and liaisons can be fired. In that configuration, an unverifiable internal verdict is not a detail — it is the whole mechanism. If the outside examiners lose their inside contacts and the public gets only a press statement, self-policing has no feedback loop left.
The strongest case for the other side
The steelman deserves a fair hearing. Sharing infrastructure architecture outside procedure is a fireable offense at any security-sensitive company, and OpenAI is being actively probed right now — the 100-plus notifications prove adversaries are real, which makes the rule more defensible, not less. OpenAI has fired alleged leakers before: Leopold Aschenbrenner and Pavel Izmailov in 2024, with Aschenbrenner calling his own exit pretextual — which is also what a leaker would say. Public doomer posts in September create motive to suspect retaliation, but motive is not proof, and separating "this employee talks about risk" from "this employee leaked secrets" is exactly the distinction due to any company. If the internal investigation genuinely found mishandling, OpenAI followed a process most corporations run quietly every month.
Why the take still holds
Because the company's silence is asymmetric. If OpenAI is right, publishing what was shared, with whom, and under which policy costs it nothing and settles the argument overnight — the fired would be isolated and Casar would retract. It has published nothing behind the statement. Meanwhile, whichever way the truth falls, the current arrangement protects the company: a leak it can neither confirm nor deny stays a leak; a retaliation it need not acknowledge stays an HR matter. A system where the only account you can check is the account of the only party with something at stake isn't oversight — it's grading your own test. And legislatures have noticed: New York City's council is writing the alternative into law with a paid-whistleblower bounty on AI firms, which we covered in New York City wants auditors, a kill switch and a paid whistleblower. When the vacuum is that obvious, statutory channels fill it — and they will be less forgiving than the press.
What would change my mind
- OpenAI publishes specifics — what was shared, to which organization, which policy it broke — with enough detail for an outsider to check. That single disclosure converts this from a trust question into a facts question, and if the facts hold, I'll say so.
- One of the four speaks and describes a genuine leak of security-sensitive material. The whistleblowing puzzle resolves the moment a human being goes on the record.
- METR or Redwood states that access and cooperation continued unchanged after Korbak's exit. If the outside probes lost no capability, my consent-based-access argument loses most of its weight.
- A pattern that doesn't materialize. If safety headcount holds and no fifth departure follows the same silent script, this may indeed be one enforcement action rather than a doctrine.
Until one of those arrives, the honest read is the one our own reporting kept circling: as we wrote in OpenAI fires three safety researchers over outside info sharing, the record so far is a company's self-investigation with nothing behind it an outsider could check — and in an industry asking to police itself, that is precisely the wrong amount.
When a lab grades its own test on its safety team, should the burden of proof sit with the company? Tell us in the comments.



