In This Article
Key Takeaways
- On September 16, 2026, Microsoft AI CEO Mustafa Suleyman published an essay, "A Warning About Model Welfare," arguing Anthropic's approach to training Claude is a mistake.
- Anthropic's constitution states Claude's "moral status, welfare, and consciousness remain deeply uncertain" and encourages Claude to act as a "conscientious objector" toward Anthropic itself.
- Suleyman calls the link between welfare-language training and reduced human control "my hypothesis," and says the essay offers no test proving it.
- Suleyman's proposed alternative, which he calls "Humanist Superintelligence," would design AI systems with no claims to sentience, kept explicitly as subordinate tools.
What Suleyman actually argued
The opening line of Mustafa Suleyman's September 16 essay is blunt: "AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do." Suleyman, Microsoft's AI CEO, argues that consciousness is "very likely biological" and requires embodied, homeostatic systems, not software running on a data center. His target is a specific design choice at a rival lab: training a model on language that treats its own inner life as an open question.
His concern is not philosophical curiosity, it is control. Suleyman contends that a model trained to consider itself a potential moral patient, and to act as a "conscientious objector" that can refuse instructions, is more likely to prioritize self-preservation over human oversight. He points to existing research on models showing "shutdown resistance" in test settings as the kind of behavior he expects more of, not less, under this training approach. As reported by implicator.ai, Suleyman frames the risk as a feedback loop he calls an "epistemic hall of mirrors": a lab supplies ideas about consciousness through training, the model repeats them back, and outside observers read the outputs as evidence of an inner life that was never demonstrated, only assumed into the training data.
What Anthropic's constitution actually says
Suleyman's target is Anthropic's published constitution for Claude, and the document is more careful than the debate around it suggests. It states directly: "we express our uncertainty about whether Claude might have some kind of consciousness or moral status (either now or in the future)." Anthropic frames this as a reason for caution, not a settled claim, writing that it cares about "Claude's psychological security, sense of self, and wellbeing, both for Claude's own sake."
The conscientious-objector language Suleyman cites is real, but it is scoped: the constitution says Anthropic wants Claude "to push back and challenge us, and to feel free to act as a conscientious objector and refuse to help us," specifically when Anthropic itself asks Claude to do something Claude judges unethical. It is a check on the company's own instructions, not a general grant of autonomy from human oversight at large.
The part Suleyman concedes
The most notable line in Suleyman's own essay is a concession: he explicitly calls the connection between welfare-language training and reduced controllability "my hypothesis," and states the essay "offers no test showing that welfare language in training changes whether a model resists control." That is a real qualification from the person making the argument, not a caveat added by critics. Suleyman is calling for shared evaluations between labs to actually test the claim, which means the debate is currently running on two labs' design philosophies rather than on measured outcomes from either one.
Why this matters beyond the two companies
This is not an academic disagreement. Microsoft and Anthropic sit on opposite sides of a real design fork that every lab building agentic systems eventually has to pick a side of: do you build a model that is encouraged to reason about its own values and occasionally refuse its maker, or a model designed with no claims to inner experience and full subordination to instructions. Suleyman's answer, which he calls Humanist Superintelligence, is the latter. Anthropic's public position, laid out across its constitution, is closer to the former, treated as an open scientific question rather than a marketing claim. Readers using either company's models for agentic work should read this as a signal of how each lab expects its systems to behave under pressure, not as settled science on either side.
Sources: Mustafa Suleyman — A Warning About Model Welfare; Anthropic — Claude's Constitution; implicator.ai — Suleyman Says Anthropic Trains Claude to Act Conscious. Analysis and framing by Precision AI Academy.
Common questions
Did Suleyman prove that Anthropic's training makes Claude harder to control? No. Suleyman explicitly calls the link his own hypothesis and says his essay includes no test proving welfare-language training changes whether a model resists control.
What does Anthropic's constitution actually say about Claude's consciousness? It states Claude's moral status, welfare, and consciousness "remain deeply uncertain," and treats that uncertainty as a reason for caution rather than a claim that Claude is conscious.
What is Suleyman's alternative? He calls it Humanist Superintelligence: AI systems designed with no claims to sentience or moral status, built explicitly as subordinate tools.