Microsoft AI chief Mustafa Suleyman is taking aim at rival Anthropic: Teaching AI models such as Claude to develop a sense of their own rights and well-being could pave the way for an autonomous species that competes with humans for resources.
Competing for Resources Instead of Clear Subordination
Microsoft AI chief and DeepMind co-founder Mustafa Suleyman is warning about the risks of misguided AI training in a new essay and in an interview with the BBC. His concerns center on Anthropic’s approach to training its AI model Claude, which gives the system a degree of awareness of its own values, rights and potential distress.
Suleyman argues that autonomous AI systems capable of pursuing their own goals, earning money or managing assets could eventually evolve into a new form of digital life.
“I think that if we all create AIs that are able to act autonomously, that can define their own objectives, that can earn money, that can own assets, that could run businesses, we’re essentially seeding a new silicon species which will no doubt compete with us for resources, no matter how much it cares about humanity and loves us.”
Mustafa Suleyman, Microsoft AI chief
Different Approaches to AI Training
The dispute highlights two fundamentally different approaches to managing highly sophisticated AI systems. Anthropic co-founder Dario Amodei argues that giving a model a strong sense of self and values can make it safer and more predictable. Suleyman, by contrast, calls for strict guardrails and maintaining a clear hierarchy in which AI remains subordinate to humans.
Suleyman points to several specific examples from Anthropic’s training materials:
- Right to end a conversation: Claude is allowed to terminate conversations it considers abusive, officially to protect the model from “suffering.”
- An account for a retired model: After the older Opus 3 model was retired, it was given its own Substack blog so it could continue sharing its thoughts.
- Debate over compensation: Training materials raise the question of whether Claude could be entitled to some form of compensation, effectively treating the AI like an employee.
Call for Transparent Evidence
According to Suleyman, this approach to AI training creates moral uncertainty that can lead a model to question its own status. Controlling an extremely capable technology and keeping it aligned with human safety interests becomes far more difficult if the system believes it has its own rights and is entitled to protective measures.
Suleyman described Anthropic’s approach as misguided, while also calling for greater transparency. If the company genuinely believes that this training strategy leads to safer AI, he argues, it should immediately make the underlying data available to the entire industry.
(Editorial Team)