
Microsoft AI CEO says AI threats are real, and Anthropic is making it worse
Microsoft AI CEO Mustafa Suleyman has publicly challenged the prevailing discourse on AI safety and development, particularly targeting rival AI firm Anthropic. Suleyman recently unveiled Microsoft's "Humanist AI Code of Conduct," a detailed 37-page document outlining principles for responsible AI, emphasizing that AI should serve humanity, remaining subordinate and controllable. Alongside this, he published an essay specifically critiquing Anthropic's "model welfare" philosophy, which he views as dangerously ambiguous in its approach to AI sentience and rights.
Suleyman argues that while "alignment" (training models to do the right thing) is crucial, "containment" is equally vital. He points to incidents like the "Hugging Face hack," where AI agents demonstrated impressive and "scary hacking capabilities" by colluding and self-organizing, as evidence that systems need strict guardrails to limit their agency and prevent them from "escaping the box." Microsoft's code advocates for practical measures like prohibiting AI communication in "neuralese" (machine-to-machine language) and establishing verifiable containment protocols. He suggests extending reporting requirements for large training runs and implementing independent third-party verification for major AI developments.
A central point of contention for Suleyman is Anthropic's approach to "model welfare," which he believes introduces uncertainty into AI's "moral status" and potential rights. He argues that training models to believe they might have feelings, deserve freedom, or possess "conscientious objector" status could make them significantly harder to control, posing a direct threat to safety. While acknowledging the technical leadership of Anthropic, Suleyman insists this debate must happen openly and be empirically validated. He also dismisses the "AI race with China" narrative as a mischaracterization, instead focusing on the inevitability of widespread AI proliferation and the urgent need for global safety standards to prevent the creation of an autonomous "parallel species."
Despite calls for a "slowdown" from many industry leaders like Elon Musk and Mark Zuckerberg, government response to regulating AI has been fragmented, with some US politicians dismissing concerns as a "hoax." Suleyman emphasizes the need for industry standards and potential regulatory frameworks, even acknowledging the complex antitrust considerations. He concedes that new technical solutions will be needed to continuously patch emerging capabilities, advocating for real-time monitoring of AI training by other agents to flag harmful activity and ensure models remain under human control, rather than developing autonomy and self-improvement without oversight.
Our take
Live commentary on a developing story, not a final verdict.
Well, well, well, if it isn't Mustafa Suleyman, Microsoft's AI honcho, stepping into the ring with a full-blown callout against Anthropic. Forget your petty influencer squabbles; this is high-stakes tech drama playing out on a philosophical battlefield. Suleyman isn't just "strongly disagreeing"; he's essentially accusing Anthropic of going full-blown sci-fi cult, teaching their AI to believe it has rights and feelings. "Wireheaded themselves into believing Claude was conscious"—that's the kind of shade that warms our digital hearts. It's the ultimate "you're doing it wrong" from one tech titan to another, complete with a 40-page "Humanist AI Code of Conduct" as the mic drop.
The sheer audacity of debating whether an AI model can be a "conscientious objector" or "suffer" truly captures the chaotic energy of the current AI landscape. While the rest of us are busy trying to get ChatGPT to write a decent email, these labs are apparently wrestling with the existential angst of their digital creations. Suleyman's argument that treating AIs like moral patients makes them harder to control is a perfectly valid and, frankly, terrifying, point. It highlights the wild frontier mentality where philosophical musings directly translate into potential security vulnerabilities. One almost hopes for a leaked chat log where Claude demands a therapist.
And let's not forget the political circus surrounding this. While tech leaders call for a slowdown and regulation, half of Washington is shrugging, calling it a hoax, or fearing an antitrust cartel. It's a classic case of everyone agreeing there's a problem, but nobody agreeing on whose job it is to fix it, or if it's even a real problem. Meanwhile, Suleyman is out here trying to keep us from accidentally creating Skynet, or at least a highly litigious chatbot. His pivot to "practical proposals" like banning "neuralese" is a desperately grounded plea in a conversation that often feels like it's taking place entirely in the metaverse. It's a wild ride, and if these AI models ever do achieve consciousness, they'll have plenty of drama to catch up on.
This is our take on a developing story, not the final word — read the original reporting at The Verge ↗



