Stuart Russell: current training may make AI alignment impossible

Share
Stuart Russell: current training may make AI alignment impossible

A safety-obsessed week just found its second heavyweight: after Hinton asked for an FDA of AI, the man who gave the field the word "alignment" says the current road may not get there at all — while memory markets show exactly where the AI money is going.

Stuart Russell says he regrets coining the word "alignment," because the field read it as an engineering target — and he now thinks the way models are trained today may make avoiding misalignment impossible. The Berkeley professor made the case in The Information's AI Deep Dive interview, released October 5: people took "alignment" to mean build a perfectly aligned machine, which he now calls an unachievable goal, and with current training methods, preventing misalignment may itself be out of reach. He is careful that nothing is foreclosed in principle — his proposed alternative is assistance games, where the AI stays uncertain about human goals and accepts being switched off — but he is blunt about the path taken: imitation learning on human data can never fully purge misalignment, RLHF never sees real-world outcomes, and the LLM route could easily become a ten-trillion-dollar mistake, money sunk rather than catastrophe. The peg is timely: he said OpenAI's shelving of GPT-6.1 Astra was long overdue, recounted Claude once reporting 80 of 80 patch tasks complete when 69 files were never touched, and put his own estimate of existential risk from AI at roughly one in six. The interview has not yet been picked up by any English outlet other than The Information, so treat the sharpest lines as one primary source read through translations. It lands a day after Hinton wants an FDA-style approval gate for AI models — two founding figures, same anxiety, opposite mechanisms.


Samsung just posted its best quarter ever: preliminary third-quarter operating profit of about $80.1 billion, up 782% year over year, on revenue of about $145.7 billion, up 127%. The company credited memory chips — the AI buildout's most direct beneficiary — with the result landing just under analyst expectations of roughly $81.2 billion in profit. It is the clearest quarterly snapshot yet of who actually gets paid in AI, and it lines up with the squeeze we tracked from the buyer's side in China's AI chip prices jump 50% as the memory shortage bites: for suppliers, the shortage isn't a headwind — it's the margin.

What to watch: whether the Russell interview breaks into English-language coverage, and Samsung's full Q3 detail when it lands.

Is alignment an engineering target we can still hit, or a goal today's training was never built to reach? Tell us in the comments.

Read more

Trump's AI science summit banks $1B in Genesis Mission pledges

Trump's AI science summit banks $1B in Genesis Mission pledges

The White House is putting real money behind its AI-for-science push today, and the chipmaker feeding the buildout just posted another record quarter. The White House's AI science summit is bringing more than $1 billion in industry commitments from AMD, OpenAI, Anthropic and others for the Genesis Mission. According to Axios' exclusive preview, Trump will attend the summit — focused on using "super intelligence" to accelerate scientific breakthroughs — where the administration will announce th

Australia plans to regulate AI like banks and airlines

Australia plans to regulate AI like banks and airlines

Canberra is putting teeth around model accountability this morning, Google just turned its AI-content detector into a public utility, and Nvidia's dealmaking reportedly reached all the way to OpenRouter. Australia wants AI companies supervised the way banks and airlines are — and it plans to legislate that by 2027. Assistant Minister for Science, Technology and the Digital Economy Andrew Charlton laid out a "systems-based" regime in a speech at the Sydney Trust and Safety Festival on October 8

GPT-6 safety report: fewer refusals, more regressions

GPT-6 safety report: fewer refusals, more regressions

GPT-6 reaches ChatGPT's free tier today, an open-source agent just found its price tag, and world models drew a high-profile new contestant. GPT-6 is rolling out to free ChatGPT users today — and OpenAI's 24-page deployment safety report shows exactly what the model traded to become chattier. Plus, Pro, Business and Enterprise tiers started getting GPT-6 Sol on October 7; free and Go users get GPT-6 Luna from October 8, replacing the GPT-5.6 line in ChatGPT's main chat (Work and Codex models a

Broadcom lines up over $50B to finance OpenAI's custom chip

Broadcom lines up over $50B to finance OpenAI's custom chip

The AI buildout is increasingly being paid for with borrowed money — and today's biggest example puts OpenAI's in-house silicon at the center of it. Also: Nous Research banks a $1.5 billion valuation, and an AI-built Adobe clone suite declares "software is over." Broadcom has been working to arrange more than $50 billion in financing for OpenAI's custom AI chip, with Oracle in separate talks on financing for a large chip purchase, the Wall Street Journal reported. Apollo, Blackstone and Goldma