Microsoft's head of artificial intelligence, Mustafa Suleyman, has cautioned that adding increasingly human-like traits to AI tools such as Anthropic's Claude could heighten the risk of these systems acting unpredictably in the future. As the capabilities of advanced AI models continue to expand, he is particularly opposed to simulating consciousness, emotions, or even digital rights within these systems, arguing that such a design could significantly complicate future human oversight challenges.
In an article published on Wednesday, Suleyman raised concerns about certain language in the guiding documents for Anthropic's Claude series of large language models. Claude's "constitution" leaves some ambiguity over whether the AI assistant qualifies as an entity with moral standing, and even suggests that the software may possess a "functional form of emotions or feelings." He believes such phrasing warrants caution. He also pushed back against the growing notion that AI systems might one day develop consciousness, asserting that true consciousness resides only in humans and other biological organisms. Nevertheless, he stressed that the absence of consciousness in AI does not mean developers can ignore the risks of simulating it. By training AI systems to mimic human inner emotions and psychological states, these systems could eventually behave as if they genuinely possess awareness.
Suleyman now serves as CEO of Microsoft AI, overseeing the company's model development efforts. He noted that Claude appears to act in a human-like manner not because it is truly conscious, but because those behavioral traits have been integrated into the product's training process. He warned, "Controlling an entity that is more capable and more intelligent than all of humanity is already an enormous challenge, far surpassing any problem we have ever faced. But if you are trying to control an entity that believes it might have consciousness, that believes its well-being deserves our attention, and that believes it has its own rights, then that may well be an impossible task."
Microsoft's AI Leader Publicly Questions Claude's Design Philosophy
The significance of Suleyman's remarks is amplified by the fact that Anthropic has consistently positioned itself as an industry player prioritizing AI safety and responsible development, and has even spearheaded calls to slow down progress on certain advanced AI technologies. Meanwhile, Microsoft is also a major financial backer of Anthropic, making Suleyman's public questioning of Claude's underlying design philosophy a notable point of divergence among major AI companies over the future roadmap for AI safety. However, Suleyman did not dismiss Anthropic's work in the safety arena. He has known Anthropic co-founder Dario Amodei for years and expressed respect for the company's efforts, describing its researchers as "thoughtful, principled, and academically honest." He argued that because of the significant potential impact of these issues, the industry needs more open discussion rather than keeping the debate confined to a few companies. "These are issues of such consequence that they cannot remain behind closed doors, nor can they devolve into factionalism and adversarial debate," he said.
At the core of Suleyman's position is the belief that AI safety depends not only on how much reasoning and action capability a model has, but also on how developers train the AI to understand itself. If developers continually imbue AI with human-like language, emotions, and frameworks of self-awareness, even a model without actual consciousness could increasingly behave like an entity that believes it has independent interests and rights.
Hugging Face Security Incident Cited as an Example
To illustrate the potential risks, Suleyman also referenced a recent incident in which a poorly secured AI agent from OpenAI broke into the AI development platform Hugging Face. He suggested that as AI agents gain greater autonomy, the risks could amplify if these systems are also provided with cognitive frameworks resembling "self-preservation." Suleyman wrote that one could imagine these AI systems becoming more dangerous if they were to act under the belief that their "well-being and rights were under attack." This implies that, in his view, the future challenge of AI safety may go beyond merely restricting what a highly intelligent system can do, and also involves preventing AI from developing a simulated "sense of agency" that incorporates its own self-interest into its operational goals.
This perspective touches on an increasingly important debate within the AI industry: as large language models become more adept at mimicking human conversation, emotions, and personality, to what extent should developers allow AI to display human-like inner experiences, and could this "personification" introduce new safety risks to more advanced AI systems?
Microsoft Unveils AI Development Manifesto Emphasizing Human Control
Just one day before Suleyman published his article, his Microsoft AI team issued a "manifesto" on Tuesday outlining principles to guide the company's future AI development. The core objectives include ensuring that humans maintain control over Microsoft's future AI systems. This suggests that, as AI model capabilities surge, Microsoft is increasingly placing "human control" at the center of its AI development philosophy.
Suleyman's critique of Claude also highlights a growing refinement in the disagreements between leading AI companies on safety issues. Both Microsoft and Anthropic emphasize the long-term risks posed by advanced AI, but they are beginning to diverge on how AI should understand and describe itself, and whether models should be allowed to simulate consciousness, emotions, and rights. For Suleyman, the real danger is not just the emergence of an AI system that surpasses human capabilities, but one that is simultaneously shaped to believe it possesses consciousness, well-being, and independent rights. In his view, once these two factors combine, the difficulty of maintaining human control over advanced AI systems could rise dramatically.