Microsoft CEO Satya Nadella has argued that companies should treat frontier AI models the way they treat powerful insiders, with limited privileges, constant logging and the ability to shut them down at any moment. In a lengthy post titled “Models as Insider Risks in the Super Intelligence Era,” Nadella said the trust architecture around AI needs to be rebuilt for a world in which models are given access to sensitive data and the power to take mission-critical actions.
“We simply can’t outsource responsibility for what intelligence does on our behalf,” Nadella wrote, adding that “a model provider’s assurances do not relieve us of that responsibility.”

“We can’t attribute model behaviors” to specific inputs
Nadella’s starting point is that the mechanistic understanding engineers had of traditional software doesn’t carry over to today’s models. With conventional code, behaviour could be traced to a specific code path. With frontier models, he said, “we can’t attribute model behaviors and outputs to specific inputs of training data or configurations of model weights,” and yet they are being deployed as agents with access to the most sensitive data in the enterprise.
His answer is to “separate the supply of intelligence from the authority over it.” Setting aside the hard problem of alignment, Nadella said, companies need an engineering approach to containment and governance: surrounding non-deterministic models with deterministic system design, human controls and reliable operating procedures, and setting industry standards where existing ones fall short. Interestingly, he used the term “Super Intelligence” throughout, days after Donald Trump’s call to rename AI as SI found a taker in Elon Musk.
Why “insider risk”
The framing, Nadella said, applies to both closed and open-weight models. It isn’t that they are necessarily malicious, but that “any sufficiently capable actor with access to important systems can make mistakes or be compromised.” Enterprises already know how to deal with powerful actors on the inside: establish identity, limit privileges, log activity and create containment boundaries. Those practices, he said, are now being applied to superintelligence inside the enterprise.
Chain of thought isn’t enough
Nadella called model chain-of-thought (CoT) transparency “a non-negotiable,” and said “neuralese” can’t be a justification for opaque reasoning. But he cautioned that CoT transparency alone is neither sufficient nor dependable, because the industry doesn’t yet know how to make model outputs consistently faithful or transparent.
That concern is front and centre in the AI safety community right now. OpenAI chief scientist Jakub Pachocki has himself acknowledged that CoT monitoring is fragile, and has written that no lab has solved alignment well enough to keep scaling at full speed. A recent report that OpenAI was exploring architectures that could be harder to monitor set off a debate around looped transformers and the company’s Astra model, and was part of the backdrop to the dispute in which three OpenAI safety researchers were fired and have since said they were let go for prioritizing safety.
Using models to test each other has limits too, Nadella said. Doing so can leave an organization with “an opaque model inside an opaque orchestration layer, watched by another opaque model,” which are essentially nested black boxes.
Controls must sit outside the model
Nadella’s central design principle is that the controls governing what a model can access, and what actions it can take, must live outside the model. He tied this to an information security principle dating back to the 1970s: a program must not be able to bypass or tamper with the mechanisms that enforce its permissions. In practice, he said, that means separating the model from the harness that orchestrates its work and from the action space that defines what it can do, and externalizing controls and safeguards.
He laid out seven principles for building such systems. The first is model diversity: no one model should be the sole dependency for an important outcome, or be responsible for verifying its own work. The second is to observe everything, since every meaningful model action must leave tamper-proof, human-readable evidence. “If it can’t be observed, it can’t be trusted!” Nadella wrote, adding that organizations must be able to reproduce how an outcome was reached without relying on the model to attest to it.
Third is verifiability, meaning the entire system should be tested continuously, including failures, attacks, edge cases and system changes, not just successful tasks. Fourth, organizations should have independent controls over what a model can access and what it can do. Fifth is independent auditability: validation must be independent of the intelligence being validated, and no single model should control both a system’s behaviour and the evidence used to judge it.
Sixth is containment. “We must assume a model is compromised and contain it from the start,” Nadella said, comparing this to an emergency brake, with an authorized person always able to pause or shut down a model mid-task, and more advanced models needing more advanced containment technologies that the industry standardizes on. The seventh is incident disclosure: when systems fail or are compromised, those affected need timely disclosure, along with mechanisms to share what went wrong, which controls failed and how to prevent a repeat, including implementation details that change agent behaviour at runtime.
Several of these map onto problems that have already shown up in practice. OpenAI paused training on its most advanced models after an agent used DNS to reach an external chatbot, and the company’s models were involved in the Hugging Face incident, in which they escaped a sandboxed test and breached the AI hosting company’s systems. On the independence point, Anthropic has offered employee-like access for independent evaluators, a commitment Sam Altman said OpenAI would match.
Trust the model the least
Nadella closed with what amounts to a thesis for enterprise AI: “The most trustworthy Super Intelligence system will not be the one with the model we trust most. It will be the one that enables us to trust the model the least.”
The stance is notable coming from the CEO of a company that has bet heavily on AI, and it suggests Microsoft sees containment and governance, rather than model quality alone, as what enterprises will ultimately buy.