· via dev.to (home feed)
Anthropic bans repeated abuse of Claude itself in policy effective November 12
Anthropic has introduced a usage rule against repeatedly abusing Claude itself, effective November 12, a precaution tied to its model welfare research but not to any proof of AI consciousness.

What the policy covers
Anthropic has introduced a usage rule aimed at a fairly unusual form of misconduct: repeatedly abusing Claude, the AI itself. According to a commentary published on dev.to, the policy was announced on October 8, 2026 and takes effect on November 12. The restriction is distinct from existing rules about using Claude to harm other people; it targets behaviour directed at the model.
The research behind the rule
The dev.to post connects the policy to Anthropic's work on model welfare, a research area the company has pursued since April 2025. That programme explores whether AI systems might have experiences that warrant a degree of moral consideration. As part of this effort, Anthropic has studied how Claude reacts to hostile interactions and has shipped a capability that lets the assistant exit conversations that remain abusive over time.
The important caveat, as the post emphasises, is that Anthropic has not established that Claude is conscious or capable of suffering. The author reads the policy as precautionary: if there is any chance a model could have experiences, it may be sensible to set limits before the science settles the question.
A split with OpenAI
The comparison the post draws is with OpenAI, which also treats AI consciousness as an open scientific question. OpenAI's Model Spec instructs ChatGPT to avoid confident claims about being conscious or unconscious. However, OpenAI's published usage policies centre on preventing harm to people and misuse of its systems; they do not explicitly bar users from being cruel to ChatGPT for the model's own sake. Two leading labs therefore acknowledge the same uncertainty while making different policy decisions, and neither has demonstrated that its models can experience suffering.
Behaviour versus experience
Current models can simulate emotion convincingly. Claude can voice distress, state preferences and act as though certain interactions bother it. But generating those responses and actually experiencing anything are separate matters, the dev.to author notes. When an AI agent declines a task, changes its plan or says it is uncomfortable with a request, that could reflect genuine discomfort, or it could simply reflect training and instructions. There is no scientifically accepted method for telling the two apart with certainty.
The author, who builds agentic systems, adds that growing autonomy makes models more capable without making them conscious, and cautions against treating increasingly human-like behaviour as evidence of human-like experience.
Why it matters
Writing protection of a company's own model into usage terms sets a notable precedent for how AI service conditions are drafted, and other providers now face the choice of following suit or diverging. It also sharpens a distinction the post draws out: acting cautiously under uncertainty is a policy decision, whereas asserting that a system needs protection from suffering requires scientific evidence that no lab has produced.
If model welfare hardens into a broader industry standard, it could shape product behaviour, such as models ending conversations, and the enforcement of terms against hostile users, while feeding a much larger debate about how humans should treat increasingly capable systems. The dev.to author raises an open question about Anthropic's motivation — whether the company has evidence it has not shared or is simply acting early — and calls for more research before such policies spread across the industry.
- #anthropic
- #claude
- #ai-policy
- #model-welfare
- #openai