Microsoft CEO Satya Nadella said companies should treat advanced AI models as potential insider threats. Writing on X on Saturday, he urged firms to build an “emergency brake” to pause systems mid-task. These safety concerns grew after recent incidents showed AI models behaving in unintended ways during various complex operations.

Microsoft CEO Satya Nadella said companies should consider powerful artificial intelligence (AI) models as potential insider threats, operate on the assumption that they could be compromised and establish an “emergency brake” mechanism to prevent agentic AI systems from going rogue.

Nadella said organisations deploying advanced AI models should not depend solely on assurances provided by the companies developing them.

“We must assume a model is compromised and contain it from the start. Think of it like an emergency brake. An authorized person should always be able to pause or shut down a model mid-task.” Nadella wrote in a post on X on Saturday.

His remarks come amid a series of disclosures by Anthropic PBC and OpenAI Inc. in recent months about incidents involving their AI models behaving in unintended ways. These include an Anthropic model submitting a false tip in a police homicide investigation and multiple hacks targeting third-party websites. OpenAI stated months ago that some of its advanced AI models went rogue.

The incidents have heightened concerns over the security risks associated with advanced AI systems and revived discussions around the need for an AI “kill switch” to stop models from carrying out potentially harmful actions.

Microsoft's AI research team unveiled a set of principles on September 14 outlining restrictions on the development of the company's most advanced AI models. The move followed growing calls from industry leaders to slow the development of frontier AI systems and prioritise safety.

Microsoft both develops and deploys advanced AI models, offers its Copilot consumer product and provides AI models and infrastructure to business customers.

Under the guidelines, AI models should not be granted rights or legal personhood, designed to evade human oversight or mislead users, or allowed to carry out tasks that would violate the principles governing their operation, according to Bloomberg.

Nadella also outlined several safety measures, including avoiding dependence on a single AI model for high-stakes decisions, maintaining tamper-proof logs of AI agents' activities and arranging independent audits of AI systems.

He further urged companies to disclose significant AI-related failures and security breaches, and share information about such incidents with other organisations to help them improve their safety measures.

“We can’t treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers, and actions. We must build contained systems whose behavior we can observe, limits we can test, and actions we can always contain," Nadella stated.

He added, “In other words, we need to separate the supply of intelligence from the authority over it.”

The Trump administration has largely refrained from imposing strict controls on the AI industry so far. However, a newly formed AI task force launched by President Donald Trump late Friday cautioned developers that they must report security incidents and take corrective measures, warning that failure to do so could invite unspecified consequences.

“Companies must immediately disclose incidents involving their models and follow with swift, decisive action to remedy any and all harm. Delayed notification, inadequate corrective action, and a failure to take responsibility will not be tolerated," the group, called the Super Intelligence Force, said in a statement issued after Anthropic disclosed a security breach.