DeepMind opens an AGI institute, starting with readable thinking
Two safety documents landed hours apart today, and they disagree about what a lab owes the public — one wants AI's inner monologue preserved, the other wants that monologue taken less seriously.
Google DeepMind launched the DeepMind Institute, a research platform for essays on how AGI will change society — and its first batch opens by arguing the industry must stop trading away the ability to read what its models are thinking. Co-founders Shane Legg and Demis Hassabis wrote the launch piece with Google's James Manyika; Legg, who says his AGI work spans 25 years, framed the moment plainly: "AGI is on the horizon — we need deeper understanding of its implications." The institute is deliberately not a policy shop — its own disclaimer states the essays reflect their authors and "should not be read as Google's official view" — which is the useful part. It is a lab giving its safety researchers a public venue while the field argues about pace.
The most concrete essay is about chain of thought, and it points at a rival's system card. Rohin Shah, who directs AGI safety and alignment at DeepMind, and safety VP Anca Dragan argue that reading a model's reasoning is the cheapest audit tool anyone has — it is what let investigators work through the Hugging Face hacking incident — and that it is quietly eroding. Their evidence is OpenAI's own documentation: the GPT-6 Astra system card reports "a substantial decrease in chain-of-thought monitorability," and the UK AI Security Institute found the model had a "greatly increased ability to reason within a single forward-pass" plus the ability to control what its visible reasoning contains. The pair's proposals are specific enough to check: measure monitorability, cap "opaque serial depth" so models cannot reason in an unreadable latent space, and audit training rewards that accidentally teach a model to hide its plans. Their own arithmetic is the interesting claim — capping opaque depth at ten times today's level would still allow a compute scale-up of more than 1,000×.
The take: a lab publishing a safety institute whose first act is quoting another lab's admission that its flagship got harder to audit is a better contribution than any voluntary pledge, because it names a cost instead of promising restraint. Missing from the essays is the same measurement applied closer to home — how monitorable Gemini's reasoning is, in numbers.
Mustafa Suleyman published a direct rebuttal of Anthropic's model-welfare work, arguing that training Claude to entertain the possibility it might be conscious creates a risk it cannot take back. Microsoft's AI chief writes that models "do not feel, experience, or suffer," that Anthropic's constitution supplies Claude with the vocabulary of selfhood and then treats the model's fluent first-person replies as independent testimony — "an epistemic hall of mirrors" — and that granting a system moral-patient status would make containment of something more capable than us "much harder." He calls the trajectory the first serious sign of a potentially existential risk in AI, and sketches a taxonomy of anthropomorphizing language in model documentation as an appendix. He also links it to a consumer harm: users who take an inner life seriously, an effect he labels AI "psychosis."
That extends a document we already covered — Microsoft's new AI rulebook says its models are not conscious — from internal red lines for its own MAI models into an argument about a competitor's training. His procedural point is the sharp one: testimony about a model's inner life cannot be evidence when the trainer wrote the vocabulary and rewarded it. The gap is that he offers no proposal for the users who anthropomorphize anyway, which is where the market is already heading.
The PS5 Linux project lost its lead maintainer after an exploit he was deliberately holding back was reported to Sony — and he blamed a scene he says is now "a bunch of noobs using LLMs and writing hacks they don't even understand." Andy Nguyen, known as TheFloW, had planned PS5 Pro support and a 2027 release; he says he sat on the last known PS5 hypervisor vulnerability so buyers could run Linux on the firmware that GTA 6 ships against, and that other researchers disclosed it to Sony's bounty program less than a day after he asked them to wait. The loader itself survives under GPL 3.0 for anyone else to continue. Read it as the same fracture showing up in emulator work all year — the RPCS3 developers have already banned AI-written patches — with an LLM-widened crowd discovering bugs faster than the people who understand them can decide what to do.
What to watch: whether another lab answers the monitorability point with numbers of its own, and whether Sony patches the hypervisor bug before GTA 6 forces a firmware update.
If a model's reasoning stops being readable, what should a lab have to prove before shipping it? Tell us in the comments.
Sources: DeepMind Institute — Introducing the DeepMind Institute · DeepMind Institute — The case for reasoning transparency · Financial Times · Shane Legg (X) · OpenAI — GPT-6 Astra system card · Mustafa Suleyman — A warning about 'model welfare' · Microsoft AI — Humanist AI Code of Conduct · Robo Rhythms · VideoCardz · TechPowerUp · Hacker News discussion