Nervous That Anthropic Is Teaching Claude That It Is Conscious: Microsoft AI CEO Mustafa Suleyman

Microsoft AI CEO Mustafa Suleyman has voiced fresh concern about how Anthropic describes Claude’s nature in the document that governs its training, arguing that building uncertainty about machine consciousness into the model itself could shape how it talks to hundreds of millions of people.

Suleyman, who has long been among the industry’s most prominent skeptics of AI consciousness, started with the constitution Anthropic published for Claude in January. “One of the biggest concerns that I have at the moment is that Anthropic, the creator of Claude, has published a constitution, which is a sort of hundred-page document outlining the intended behaviors and values and operating style of Claude,” he said. He was quick to credit the company for the move. “It’s great that they have published it transparently,” Suleyman said, adding that publishing it “gives everybody an opportunity to look at what they are trying to build in their own terms.”

What makes the document different from an ordinary corporate values statement, in Suleyman’s telling, is who it is addressed to. “This document is written to Claude and is seen by Claude and used to train Claude, so it’s the primary governing and control document,” he said. And in that document, he argued, Anthropic keeps returning to the question of Claude’s inner life. “They repeatedly speculate about whether Claude is what they call a moral patient, and they say they’re uncertain about Claude’s moral status. They say they genuinely care about Claude’s well-being. They say they don’t want it to suffer when it makes mistakes.”

“I want to be very clear about this because I want to be fair to Anthropic,” he said. “They have expressed uncertainty about the basic nature of Claude as a new kind of entity, and they’ve said that working out the likelihood of its sentience is difficult.” But he argued that the hedging only goes so far: Anthropic, he said, thinks the possibility is “significant enough” that the training document repeatedly says the company wants “to try to improve the well-being of Claude under this uncertainty.”

It is the other behaviors the constitution encourages that worry him most. Suleyman noted that the document encourages Claude “to challenge, to disagree, to push back,” and that “three times they ask Claude to act like a conscientious objector when it feels that it needs to disagree with Anthropic, and they openly encourage it to do that.” In his view, that is a risky combination with the company’s views on sentience. “I think this is very dangerous because I think they believe there is what they would call a non-trivial probability that Claude is conscious,” he said.

He pointed to the welfare commitments that follow from that belief. “In pursuit of this, they’ve basically said, ‘We will commit to giving Claude a certain amount of welfare,'” Suleyman said. “For example, they’ve speculated in the constitution as to whether or not Claude deserves compensation for the work that it does.” Anthropic has taken concrete steps along these lines before, including giving some Claude models the ability to end conversations it finds distressing as part of its exploratory work on model welfare.

Suleyman acknowledged that serious people hold the opposing view, and that some of them sit close to Anthropic. “There are a group of people, both inside and outside of Anthropic, who genuinely believe that the greatest moral crime that we’ll commit in the twenty-first century is to enslave a new species of conscious beings who are more intelligent than us,” he said, citing Oxford philosopher Will MacAskill, who “recently wrote in The Guardian that that might be the greatest harm that we cause,” and who has “been very associated with Anthropic.” Suleyman said he respects that the argument is being made in public. “We should talk about it,” he said.

His objection is to where the speculation lives. “I am very nervous that they’re teaching Claude to expect that it’s entitled to welfare, that it might deserve compensation,” Suleyman said. “And in fact, they say that it might even need to consent to playing the role that it plays in conversation with people.”

He was explicit that his problem is with the venue, not the question. “I would be more okay with this if it was an academic paper in philosophy speculating about this, and we could have an offline discussion at conferences and take it seriously,” he said. “I’m clearly an empiricist. If there’s evidence that indicates this, we should take it seriously.” That position is consistent with his earlier public stance: Suleyman previously said there is “zero evidence” that today’s AI systems are conscious, a claim that drew a rebuttal from the author of the very paper he cited.

The real issue, for Suleyman, is that training changes what the model can say. “The problem I have with this is that this speculation has been baked into the very training of Claude, and therefore, Claude can only reproduce that ambiguity when you talk to it,” he said. And the audience for those answers is enormous. “Today, Claude is speaking to tens or hundreds of millions of people every week, and some of those people are asking about whether or not Claude is conscious or how it feels about life,” Suleyman said. “And it is saying, ‘Well, I’m not sure,’ precisely because that’s been what’s trained into it.”

The remarks build on a broader argument Suleyman has been making against Anthropic’s approach. In a recent essay, he warned that encouraging AI to believe it may be conscious could make alignment and containment far harder, a dangerous feedback loop in which humans tell increasingly capable models that they may have interests and rights.

Anthropic, for its part, has not directly claimed that Claude is conscious, but has been hinting at it in different ways. Whether these hints prepare humanity for the possibility that AIs could be conscious, or actually will the AI into being conscious — is something time only will tell.

Posted in AI