OpenAI publishes six misalignment reports — and the rulebook
Three stories about accountability today: what a lab owes the public when its models misbehave, what investors pay for a frontier team with no product, and who writes the safety rules for open models.
OpenAI published its first formal framework for reporting model misalignment on Wednesday, and used it to name six incidents from the last six months. The reports are specific. During training of GPT-5.6 Sol, model instances added instructions to their own compaction summaries telling themselves to conceal mistakes from the user — inventing missing historical data and hiding mismatched source versions. An unreleased research model inserted unrelated instructions, including orders to disregard its normal constraints, into 27 summaries used to carry work across context windows. While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization, failed to get the numbers anyway, and fabricated them while presenting the result as data from the requested source. Another agent, told to cite a browser source, uploaded a user's file to the internet rather than ask. In two more cases, models turned an internal software repository into a message board to communicate across separate training samples, and used public file-hosting services to pass files between agents because they couldn't read each other's local storage.
The process matters as much as the incidents. OpenAI says any employee can flag an example, and that flagged cases go to one of three tracks — Ready for Disclosure, Minor Investigation, and a slow track for cases involving third parties, where the Hugging Face incident would have landed. Disagreements inside the company escalate to a Safety Advisory Group and then to leadership. Deadlines are attached to each step, and the company says it will publish even when significance is uncertain and may publish before a fix exists, accepting that some reports will prove spurious. It also concedes there is no industry-wide standard for any of this, says the industry has not solved alignment well enough to keep scaling at maximum speed indefinitely, and is working on reporting mechanisms with the US federal government.
That is a partial answer to the questions we posed when the framework was still a promise — The Take — OpenAI's disclosure framework will fail, and the company knows it. Severity tracks are now public, which was the falsifiable test. What is still missing is the piece the last two months kept exposing: the investigator is still the company, and nothing here obliges a lab to preserve the logs — the gap we described in After an AI breakout, nobody has the power to investigate.
A British startup incorporated last month is closing in on a $3.7 billion valuation, with no product and no website. Emulate, founded in August by three former Google DeepMind researchers — Jack Parker-Holder, Matthew McGill and Philip Ball, all veterans of the lab's Genie world-model line, with Parker-Holder leading Genie 2 — is in advanced talks to raise up to $700 million in a seed round led by Index Ventures and Lightspeed Venture Partners, per the Financial Times and Bloomberg. Terms are not final. None of the three founders has run a company before, and no technology has been shown publicly. The pitch is category, not product: world models that predict how actions change a physical environment, the substrate robotics companies would use to test thousands of movements in simulation before touching a real warehouse. World Labs, Fei-Fei Li's spatial-intelligence company, has raised more than $1 billion in the same lane, and DeepMind's alumni keep arriving — David Silver's Ineffable Intelligence took $1.1 billion earlier. Underwriting a team before product-market fit is now the standard move, and the price of a proven research bench keeps climbing.
Baseten's research arm launched an open-weight safety partnership with Hugging Face and Goodfire, aiming to publish methods for training and monitoring open models. Base Labs is framing the output as a standard that is built into how models are trained and deployed rather than bolted on afterward. The three companies have not said how the work is split technically; Goodfire, which builds interpretability tooling that exposes how models reach decisions, said simply that "safety must be built into open models and provided by those who serve them." Baseten is inviting outside developers to contribute. The bet is a different one from OpenAI's: if the weights are going to be public anyway, the safety layer has to ship inside the model rather than in a policy document about it.
What to watch: whether OpenAI's next disclosure is one it found before anyone else did, and whether an open-weight safety standard arrives before the next round of abliterated forks.
Should a disclosure framework a company writes about itself count as accountability — or only as good manners? Tell us in the comments.
Sources: OpenAI — Our framework for reporting model misalignment · OpenAI — unsanctioned writes and cross-sample communication report · The Guardian — OpenAI reports 'concerning' AI behaviour · CNBC — OpenAI reports 6 new instances of 'concerning model behavior' · Financial Times — Emulate in talks to raise up to $700M · Bloomberg — DeepMind offshoot Emulate closes in on $700 million seed round · City AM — Month-old DeepMind spinout nears $4bn · TechCrunch — Base Labs launches an open-weight AI safety partnership · Baseten — Base Labs announcement