IT Home, October 11 — Microsoft CEO Satya Nadella believes that in an era of rapidly advancing AI, simple strategies such as self-monitoring with superintelligence are insufficient.
In a lengthy post on the evening of October 10 Beijing time, Nadella pointed out that traditional software systems have been widely deployed over the past few decades, and humans have developed tools and capabilities to track the behavior of specific code paths. However, AI models are more powerful than traditional software, yet they remain black boxes. We have deployed these complex agent systems and models, granted them access to our most sensitive data, and entrusted them with the ability to perform critical tasks on our behalf.

Nadella believes we should take a step back to reconsider the trust architecture of this new era. Model providers cannot outsource responsibility or treat superintelligence as a nested black box, simply accepting or rejecting its suggestions, answers, and actions. We must build closed systems that can observe their behavior, test their limits, and always accommodate their actions.
Nadella stated that treating both cutting-edge closed-weight and open-weight models as internal risks is one way to build such a system. This is not because AI models are inherently malicious, but because any actor with sufficient capability and access to critical systems could make mistakes or be compromised—and the architecture for containment and control must account for this.
Nadella mentioned that developers should design these systems around the principle of observability; IT Home lists the specific principles as follows:
- Model diversity: No single model should be the sole dependency for critical outcomes, nor should any model be responsible for verifying its own work.
- Observe everything: Every meaningful model action must leave behind tamper-proof, human-readable evidence. If it cannot be observed, it cannot be trusted. We must be able to reproduce how results were achieved without relying on the model to prove it.
- Verifiability: We must continuously test the entire system, including failures, attacks, edge cases, and system changes—not just successful operations.
- Independent control: The organization should be able to independently determine what the model can access and what actions it can take.
- Independent auditability: Verification must be independent of the intelligence being verified. No single model should simultaneously control the behavior of the system and the evidence required to determine whether that behavior aligns with the original intent.
- Containment: We must assume that models can be compromised and implement containment from the outset. Think of it as an emergency brake. Authorized personnel should always be able to pause or shut down a model during operation. More advanced models will require more sophisticated containment techniques, and we need to standardize these approaches.
- Incident disclosure: When these systems fail or are compromised, we must promptly disclose information to affected parties and establish mechanisms to share the causes of the failure, which controls failed, and how to prevent recurrence, as well as share lessons learned across the entire industry. This should include implementation details related to altering agent behavior during runtime.
