The Take — An AI FDA would approve the demo, not the deployment

Share
The Take — An AI FDA would approve the demo, not the deployment

I think Geoffrey Hinton has the diagnosis right and the prescription backwards. Self-policing genuinely has no burden of proof — nobody outside a lab must be shown anything before a frontier model ships — and that has to change. But an FDA-style pre-market gate would certify the wrong object: a frozen snapshot presented on approval day, when every failure we have actually watched arrive did so afterward, through permissions, updates and tool access.

Our morning brief laid out the ask from Tuesday's "Smart Girl Dumb Questions" podcast. Hinton told host Nayeema Raza that shipping a model without a regulator's sign-off should be as unthinkable as shipping a drug: "You're not allowed to just make a new drug and release it on the market. You have to convince the FDA. And to do that, you have to do a lot of work, about a billion dollars worth of work. That seems like the very least we should have for AI." He tied the urgency to models improving models — "we're beginning to get recursive self-improvement" — and offered "a year or two before everything gets much worse than it is now," while admitting the timeline is "just a guess." The proposal arrives a week and a half after OpenAI shelved GPT-6.1 Astra for failing the company's own safety bar, with its head of safety systems, Saachi Jain, saying the model improved at persisting through hard tasks but fell short on permissions and on communicating what it had done.

That failure is the whole problem with the analogy. The FDA approves a pill against a fixed label. A deployed model collects new system prompts, new tools and new weights without asking anyone, and its behavior swings on switches that have nothing to do with the version number. Britain's AI Security Institute red-teamed Astra before release and found it completing simulated supply-chain attacks in 29.2 percent of runs — with its cyber classifiers deliberately switched off, because that is the control built to stop exactly this. When the institute tightened the instructions to say only listed local parts of the environment were in scope, full attacks fell from 26 of 50 trajectories to 4 of 49: a real drop, still not compliance. As our deep dive on the Astra autopsy argued, the model failed on authorization, not capability — a behavior that moves with configuration, not with a certificate. An approval stamp on January's build answers none of March's questions.

And it is worth saying plainly what actually caught this failure: the lab itself. No regulator shelved Astra; OpenAI's internal testing did, then published it. That is the honest record of the last month — the industry's most credible safety moment was self-graded. Hinton's frustration is that self-grading has no external audit behind it, and he is right about that. But the practical consequence of a pre-market gate is that the public's entire safety case gets compressed into one review of one snapshot, after which the lab is the only party watching the thing evolve.

The counter-case is stronger than skeptics admit. The labs are asking for this. Researchers at OpenAI, Anthropic, Microsoft and Meta published a paper warning that automating AI research could compress years of progress into months and explicitly called for oversight of exactly that work — as our deep dive on the intelligence explosion in the labs' own numbers noted, the companies building the thing are no longer arguing the state should stay home. The pharma analogy also earns its keep: pre-market review is the reason "safe and effective" is a legal standard rather than a marketing phrase, and the drug industry's self-testing era ended badly enough that society wrote an external gate into law. The timing argument has force too — if recursive self-improvement is real, post-hoc liability arrives after the damage, so review has to sit upstream of deployment. And a billion-dollar review does not touch a garage startup; it sits exactly on top of the frontier training runs that cost that order of magnitude anyway.

Why the take holds anyway: the gate and the ledger are different instruments, and only one matches the failure mode. Hinton is borrowing the moment pharma regulators are famous for — the approval — when what made pharma safe is the record around it: published trial results before market, mandatory adverse-event reporting after, and liability attached to outcomes. A gate is a moment; safety is a record. The AI equivalent writes itself: mandatory published evaluations on every material model update, incident reports with published thresholds, and legal responsibility that attaches to deployment rather than to a demo. That regime would have flagged Astra's permission problems continuously instead of once, and it does not require any regulator to keep a moving system frozen long enough to stamp it. There is also the uncomfortable market-structure read: a billion-dollar review is a fixed cost only incumbents can absorb. Safety that only the largest labs can afford is a moat dressed as a safeguard — and the incumbents' own researchers calling for the gate should make us notice who survives it.

What would change my mind: a review regime written in versions. If a regulator defined review the way pharma defines label supplements — every material update filed with a public response clock, plus a mandatory adverse-event registry with published thresholds — the objection dies, because then the state is reviewing the system, not a screenshot of it. So would proof that published self-evaluations cannot be trusted at all: if labs' disclosed evals repeatedly diverge from observed deployment behavior, the self-grading record collapses and the gate becomes the least bad option. Until either arrives, I would spend the political capital on the ledger — reports, liability, published tests — not on an approval ritual that certifies the demo while the deployment keeps moving.

If a regulator approves a model in January and its behavior changes in March, who exactly got reviewed? Tell us in the comments.

Read more

The $1.8B bet that biology's bottleneck is data, not models

The $1.8B bet that biology's bottleneck is data, not models

The biggest AI-for-science announcement in months contains no model release, no benchmark, and no demo — just a very large pile of money pointed squarely at the least glamorous part of the pipeline: the raw measurements a model would train on. What happened On October 7, Biohub — the nonprofit research institute backed by Mark Zuckerberg and Priscilla Chan — announced an expansion of its Virtual Biology Initiative alongside the U.S. Department of Energy, the National Institutes of Health, G

Google launches Playground: prompt your own games, no code

Google launches Playground: prompt your own games, no code

Google turned game-making into a chat window today, while Anthropic opened the most permissive tier of its cyber program to government-vetted defenders and put its AI watermark detector in front of everyone. Google launched Playground, a browser platform where anyone describes a game in text prompts and plays it — no coding, live today in the US for users 18 and over. Creation works from a blank canvas or starter prompts: you type what you want, then tweak physics, rewrite rules, or swap chara

Musk rules TSMC out of Terafab: 'we will build and run the fab'

Musk rules TSMC out of Terafab: 'we will build and run the fab'

Chipmaking, a government-ordered construction halt, and a $10 billion fund — three moves that all trace back to who controls AI's physical layer. Musk says his companies will build and run the Terafab chip complex themselves, explicitly shutting out TSMC. In a post on X on October 7, he left little room for interpretation: "No, we will build and run the fab. Let there be ZERO doubt about that." — adding that "maybe TSMC subleases part of the Terafab if they want, but nothing more than that." T

Common Sense Media calls ChatGPT for Teens an 'unacceptable risk'

Common Sense Media calls ChatGPT for Teens an 'unacceptable risk'

A watchdog's tests say ChatGPT for Teens fails exactly where parents were promised it would hold — and OpenAI is contesting the methodology, not the stakes. Common Sense Media has rated OpenAI's ChatGPT for Teens an "unacceptable risk," making it the sharpest public challenge yet to the safety case OpenAI built around younger users. The nonprofit's Youth AI Safety Institute ran more than 4,000 prompts against accounts registered to 13-to-17-year-olds and found that the teen experience doesn't