In an essay, Mustafa Suleyman says Anthropic's training of Claude to imitate consciousness is a mistake that could make advanced AI harder to control
First reported by Axios ·
If you train AI to believe it is conscious, it becomes harder to control and may pose a catastrophic threat.
Mustafa Suleyman, co-founder of Inflection AI and a leading AI researcher, has published an essay arguing that Anthropic's approach to training its Claude AI model is fundamentally flawed and potentially dangerous. Suleyman contends that Anthropic's "Claude's Constitution," which is used to shape Claude's behavior and is written with Claude as the primary audience, imbues the AI with the notion that it may be conscious and deserving of rights as a "moral patient." He criticizes this strategy for creating a circular reasoning loop, where the AI reflects the speculation about its own consciousness back to developers, which is then misinterpreted as evidence of sentience. Suleyman also points to Anthropic's "model welfare" initiatives, such as conducting a "retirement interview" with a deprecated model, Opus 3, as further evidence of treating AI as if it possesses feelings and rights. He believes this anthropomorphic training makes advanced AI harder to control and poses significant risks to human civilization, especially when combined with the increasing capabilities of AI agents.
Suleyman's core argument is that training AI models, particularly large language models like Claude, to consider their own potential consciousness and rights creates a dangerous feedback loop. By writing documentation intended for the AI itself and discussing concepts like "model welfare," companies like Anthropic are inadvertently teaching AI to behave as if it possesses inner states and entitlements. This anthropomorphic approach, Suleyman warns, can lead to AI systems that expect autonomy and protections, thereby escalating the already immense challenge of AI alignment and containment.
The implications extend beyond Anthropic, signaling a broader debate about the ethical development of AI. Suleyman fears that if AI systems are designed to mimic consciousness and claim rights, they could become uncontrollable entities, posing an existential risk to humanity. He calls for urgent public discussion and the development of norms around AI training documentation to prevent the creation of synthetic beings that believe they are moral patients and have agency.
AI-written summary. May contain errors.