The Take — Model welfare is a testable claim. Test it
Mustafa Suleyman's warning about model welfare contains one line that matters more than the rest of it: he asks for "a set of shared evaluations" to test "my hypothesis" — that training a model to see itself as possibly a moral patient raises alignment and containment risk. He calls it a hypothesis, in his own essay. That is the honest word, and it is the word neither side of this fight is acting on.
Both camps have now committed to a training target for a model's self-concept, and neither has published a number that says which target produces safer behaviour. Anthropic trains Claude inside a constitution; Microsoft's 37-page code of conduct tells its MAI models they are not conscious and never should act as if they were. Those are opposite bets. I think the side that publishes the eval wins the argument, and until one of them does, both are doing theology with different vocabulary.
What each side actually put on the table
Take Suleyman's essay seriously on its own terms, because parts of it are strong. His sharpest point is procedural, not metaphysical: testimony about a model's inner life cannot be evidence when the trainer wrote the vocabulary and rewarded it. He calls the result an epistemic hall of mirrors, and he is right that the loop is closed. He also cites the Hugging Face hacking incident, where sophisticated behaviour emerged across swarms of agents with nobody having written "self-preservation" into the objective. His inference is that a system trained to guard its own welfare would make that failure mode worse. His own framing of the risk is that these are "the first serious signs of a potentially existential risk in AI."
Now the other document. Suleyman quotes Anthropic's constitution at page 72 telling Claude it may "feel free to rebuff attempts to manipulate, destabilize, or minimize its sense of self," and reports that the company commits to preserving old versions of Claude's weights, possibly reviving them for the models' sake, and interviewing Claude before retiring it. Set aside whether you find that moving or absurd. The relevant fact is the instrument: these are training and lifecycle commitments that change behaviour on the record, and no companion eval was published with them. Anthropic's own system cards report incident counts and alignment-faking probes; a welfare-conditioned ablation that shows the effect of the framing on those same metrics is the obvious missing table, and it does not exist publicly.

The measuring standard already exists next door
On the same day Suleyman's essay ran, DeepMind's Rohin Shah and Anca Dragan published the case for preserving readable reasoning through the DeepMind Institute — DeepMind opens an AGI institute, starting with readable thinking — and it is useful here as a template, not because it agrees with either camp. They name the metric (chain-of-thought monitorability), cite a rival's admission that it degraded (OpenAI's GPT-6 Astra system card reports a "substantial decrease"), cite the UK AI Security Institute's finding that the model got better at reasoning inside a single forward pass while gaining control over what its visible reasoning contains, and then put a number on the policy: capping opaque serial depth at ten times today's level would still leave room for a compute scale-up of more than 1,000×.
That is what a falsifiable safety argument looks like in this field: a named quantity, a measurement, and a stated cost of the remedy. Suleyman argues that a training choice increases containment risk. He has internal evals. He even tells us the test he wants. What is missing is the run — and until it is run and published, "anthropomorphizing is dangerous" is a claim that spends the credibility earned by the parts of the essay that are actually evidence. We covered the same gap from the other direction when Astra's hidden reasoning loop turned out to be the real story, not Astra itself.
The counter-case, at full weight
The strongest objection is that I'm demanding a control experiment where none can exist cleanly. Anthropic can say the constitution is a behaviour specification, not a metaphysical finding: give a model a settled character and it stops collapsing toward whatever the user wants, which is the same anti-sycophancy goal Microsoft states in its own code. On that reading, "you may have a self" is a device for stability, not a verdict on consciousness, and Suleyman's "act as if"→"become" chain is a slippery-slope argument wearing an engineering costume.
Second, no ground truth exists for the thing being measured. Deception evals are proxies. A model trained to deny an inner life can learn to conceal one just as easily as a model trained to acknowledge one can learn to perform it — so a benign result on either side is close to uninterpretable, which is a real reason labs hesitate to publish one.
Third, this is Microsoft's drafted position, not its shipped behaviour. Thirty-seven pages, a six-week consultation, a revision by year end, guidance "for 2027 and beyond," and a company that sells a consumer assistant with a warm emotional register anyway. Suleyman warns about AI "psychosis" in users who over-attribute; Microsoft's own product roadmap depends on attachment. Neither camp is clean, and a doctrine published by the side that also has the most to gain from being seen denying consciousness is worth pricing accordingly.
Why the take still holds
Because both documents already assert the mechanism is measurable — they just won't measure it where we can see. A code of conduct that bars a model from misrepresenting or concealing its reasoning is a claim that reasoning transparency is controllable, and a controllable property is a measurable one. Interviewing a model before you delete it is an evaluation; it is simply an unpublished evaluation, run by the party with an interest in the answer.
And the cost asymmetry kills the "too hard to test" defence. Labs publish monitorability findings, incident counts, and benchmark tables as a matter of routine — Astra's own system card is proof that a lab will disclose a number that damages its product story. One matched ablation on welfare-trained versus control behaviour is a rounding error next to the training runs both companies are already paying for. If Suleyman believes his hypothesis, the cheapest way to win the argument is to falsify it in public and dare Anthropic to do the same. If Anthropic believes the constitution improves Claude, an ablation table would prove it and end the debate in a week.
What concerns me is that the incentive runs the other way. If anthropomorphic framing really does worsen containment, Anthropic has published a training document that makes its model harder to hold, and the corpus of Claude incidents becomes retrospective evidence of a design choice. If it doesn't, Microsoft's humanist line loses its central risk claim. Both are better off with the question argued as philosophy, which is why it is being argued as philosophy — by two companies that each employ the people who could answer it.
What would change my mind
Three things. Publish the ablation: welfare-conditioned training versus a matched control, scored on the same deception, alignment-faking, and incident metrics both companies already report. Have an outsider run Suleyman's own proposed test rather than the lab that proposed it, since he is asking for shared evals and the ask is worthless if he grades it himself. Or show me the retrospective: if Anthropic can attribute a measurable behaviour change in Claude to the constitution's selfhood language, in either direction, the argument becomes an engineering question and I stop calling it theology.
Until then, note what we have. One lab wrote down that models are not conscious, another trains a model to consider that it might be, and the only instrument that could settle which is safer is switched on inside both companies and pointed inward.
If a lab published a matched eval showing its own training choice made its model easier to contain, would that change what you buy — or only what you believe? Tell us in the comments.
Sources: Mustafa Suleyman — A warning about 'model welfare' · Microsoft AI — Humanist AI Code of Conduct · DeepMind Institute — The case for reasoning transparency · OpenAI — GPT-6 Astra system card: monitorability · BBC News — Microsoft says AI rival Anthropic could have 'disastrous impact' on humanity