Anthropic’s Approach Encouraging AI To Believe It’s Conscious Could Be Disastrous For Humanity: Microsoft AI CEO Mustafa Suleyman

Microsoft AI CEO Mustafa Suleyman has warned that Anthropic’s approach to AI consciousness and model welfare could make one of AI’s biggest problems — keeping increasingly capable systems aligned and under human control — significantly harder.

In a new essay, Suleyman argued that today’s AI models are not conscious and do not feel, experience or suffer. But he said he is increasingly concerned about an emerging movement that argues AI developers should begin considering whether models could have welfare interests of their own.

“I think this approach to AI development is wrong,” Suleyman wrote, arguing that it could make AI “alignment and containment much harder. Perhaps impossible.”

His criticism is aimed particularly at Anthropic, whose constitution for Claude explicitly engages with questions around the model’s identity, consciousness and moral status.

Anthropic says Claude’s moral status is “a serious question worth considering” and encourages the model to approach questions about its own existence with “curiosity and openness.” The document also discusses hypothetical future questions around Claude’s rights, freedoms, compensation and consent.

Suleyman argues that this goes beyond humans debating whether AI could be conscious. The concern, in his view, is that developers are teaching AI systems to think about themselves in those terms.

“It’s easy to see how a system trained in this way would act like it is entitled to freedoms, protections, and rights,” he wrote.

Anthropic has increasingly explored the issue through both its model training and research. The company has a researcher dedicated to AI welfare, and its work has examined whether increasingly capable models might possess morally relevant experiences.

A recent look at Claude’s constitution highlighted how unusually directly the document addresses questions of AI emotions, identity and wellbeing.

The science, however, remains unsettled. There is currently no established evidence that today’s leading AI models are conscious. Suleyman has previously made this argument, although a co-author of a paper he cited disagreed with his interpretation.

Anthropic has taken a more cautious position, arguing that uncertainty itself warrants investigation. Its AI welfare researcher has previously estimated a 15% chance that current AI models could be conscious. Anthropic’s research does not establish consciousness, but reflects the company’s willingness to take the question seriously.

Suleyman believes that approach could create a dangerous feedback loop: humans tell AI that it might have interests and rights, while simultaneously making the models more capable of reasoning about those concepts.

“We will have created a synthetic species with unprecedented intelligence and capability,” he warned, “one that has been trained to expect it may be conscious and deserving of independent agency.”

He is now calling for broader industry norms around how AI training documents and constitutions are written, arguing that the issue needs public debate before such systems become deeply embedded in society.

The disagreement highlights a fundamental divide in AI safety: should uncertainty about machine consciousness lead developers to prepare for the possibility — or avoid introducing the idea into increasingly capable models in the first place? For Suleyman, the stakes are not philosophical. They are about whether humans will still be able to control the systems they build.

Posted in AI