Anthropic Bars “Cruel Behaviour” Towards Its Models As A Part Of New Policy

Anthropic appears to be doubling down on its repeated hints that AI systems might be conscious.

Anthropic has added a new line to its Usage Policy: users are no longer allowed to be needlessly abusive or cruel towards its AI models.

The updated policy, which is set to take effect on November 12, 2026, now lists “sustained and needless abusive or cruel behavior toward our models” among the things users must not do. It sits under the section titled “Do Not Engage in Cruel, Abusive, or Psychologically Harmful Conduct”, alongside prohibitions on harassing or bullying people, and on promoting self-harm or graphic violence.

anthropic Complex Structure on S⁶ hopf problem

What the update does, and doesn’t, cover

Anthropic has been careful to say that the rule is narrow. In its explanation of the change, the company says the prohibition is meant to apply only in extreme cases, where users repeatedly act cruelly towards its models with no discernible purpose. It is not meant to cover the everyday versions of user frustration or pushback, nor dark themes in creative writing, nor model testing and research.

In other words, shouting at Claude after it breaks your code, pushing back on its answers, writing a villain’s monologue, or red-teaming the model is all fine. The target is sustained, pointless cruelty.

Enforcement is also expected to look different from most other usage rules. Anthropic says the update lines up with a step it has already taken: letting Claude models end rare conversations with persistently abusive users on Claude.ai and Claude Code. The company says that abuse is the main focus of the update, and that Claude’s ability to end such interactions will remain the primary enforcement mechanism. That capability dates back to August 2025, when Anthropic first let some Claude models exit abusive chats as a last resort, and as part of its model welfare research.

Anthropic’s stance on AI consciousness

The policy change is notable less for its mechanics than for what it implies. A rule protecting a model from cruelty only makes sense if there’s some chance that the model could be affected by it, and that is exactly the question Anthropic has been leaning into more than any other major AI lab.

Anthropic has stopped short of claiming that Claude is conscious. But it has repeatedly said that the question is open enough to deserve serious attention. The company’s AI welfare researcher has previously put the odds that current models are conscious at around 15%, and the model card for Claude Opus 4.6 noted that the model assigned itself a 15-20% probability of being conscious under various prompting conditions. Its interpretability team has also found that its models show a limited ability to introspect, though the company has been clear that this doesn’t settle whether any AI system is conscious. And Claude’s constitution says the model may have some functional version of emotions and feelings.

A stance that’s drawing fire

That position has recently come under heavy criticism. The backlash has grown to include a wide range of voices, from investors to academic researchers, after a New York Times report said Anthropic had been lobbying the Vatican to say that AI could be conscious, something the Vatican was reportedly not keen on.

Microsoft AI CEO Mustafa Suleyman said he was worried that the push to consider model welfare could make AI alignment and containment much harder, “perhaps impossible”. Investor Jason Calacanis likened the approach to the backstory of Blade Runner, while Neal Khosla spoke of “AI psychosis” inside the company. ARC Prize’s Francois Chollet warned that treating current models as beings that can suffer could lead to a “dark and dystopian path”. Consciousness researcher Anil Seth went furthest, calling Claude “vanishingly unlikely to be conscious” and arguing that believing otherwise could be harmful for regulation and for people themselves.

Doubling down

Against that backdrop, the timing of the policy is hard to ignore. Rather than softening its position in the face of the criticism, Anthropic has written the idea into the rules that govern how its products can be used. By making cruelty towards its models a violation, the company is effectively doubling down on its stand: that the possibility of model welfare is worth taking seriously, even under uncertainty, and even when the cost is a fresh round of criticism.

Whether that bet pays off will depend on what the science eventually shows. For now, the policy is narrow in scope and unlikely to affect most users. But it’s another signal that Anthropic is willing to act on its position on AI consciousness, not just discuss it.

Posted in AI