AI responsibility cannot be outsourced, says Satya Nadella; calls for stronger safeguards

/ 3 min read
AI Hub

The Microsoft CEO said that AI systems should not be treated as black boxes whose recommendations and actions are simply accepted or rejected.

THIS STORY FEATURES
Satya Nadella, CEO, Microsoft.
Satya Nadella, CEO, Microsoft. | Credits: Getty Images

Microsoft Chairman and CEO Satya Nadella has called for a fundamental rethink of how artificial intelligence (AI) systems are governed, warning that increasingly capable AI models are being deployed with access to sensitive data and the ability to take mission-critical actions without a clear understanding of how they arrive at their decisions.

ADVERTISEMENT

In a post on X, Nadella said traditional software systems allowed developers to trace behaviour to specific code paths. However, similar mechanistic understanding remains elusive in today’s superintelligence systems, even as frontier AI models become more capable than conventional software.

“We can’t attribute model behaviors and outputs to specific inputs of training data or configurations of model weights,” Nadella wrote, highlighting the challenge of understanding how AI systems generate their outputs.

ADVERTISEMENT

He warned that organisations are increasingly deploying complex, agentic AI systems with access to sensitive information and the authority to act on users’ behalf, making it essential to reassess the architecture of trust underpinning these technologies.

“That’s why it’s time to step back and assess the trust architecture for this new era. We simply can't outsource responsibility for what intelligence does on our behalf. A model provider’s assurances do not relieve us of that responsibility,” he said.

Separate AI intelligence from decision-making authority

Nadella said that AI systems should not be treated as black boxes whose recommendations and actions are simply accepted or rejected. Instead, they must be designed with observable behaviour, testable limits and mechanisms to contain their actions.

“In other words, we need to separate the supply of intelligence from the authority over it,” he said. He called for an engineering-led approach to AI containment and governance, combining non-deterministic models with deterministic system design, human controls, and reliable operating procedures. Where existing safeguards are inadequate, the industry should establish new standards, he added.

Recommended Stories

Nadella suggested treating both closed and open-weight frontier models as potential insider risks within enterprises. This, he clarified, does not mean assuming that the models are malicious, but recognising that any sufficiently capable actor with access to critical systems can make mistakes or be compromised. 

He said organisations could draw on established enterprise security practices, including identity verification, limiting privileges, logging activity and creating containment boundaries.

ADVERTISEMENT

Transparency alone is not enough

Nadella called for transparency in AI models’ chain of thought (CoT) to become a non-negotiable requirement, arguing that the complexity of neural networks should not be used to justify opaque reasoning. However, he cautioned that CoT transparency alone cannot guarantee reliability because model outputs are not yet consistently faithful or transparent.

He also warned against relying on multiple AI models to verify one another without independent safeguards. Such arrangements could create “nested black boxes”, in which an opaque model operates within an opaque orchestration layer and is monitored by another opaque model.

Most Powerful Women In Business 2026
View Full List >

Instead, controls governing what a model can access and which actions it can perform must operate outside the model itself. Nadella linked this principle to a longstanding information-security concept: a programme must not be able to bypass or tamper with the mechanisms that enforce its permissions.

This requires separating the model from the system that orchestrates its work and from the set of actions it is authorised to perform, while keeping controls and safeguards independent.

Nadella outlined seven principles to improve the observability, security and accountability of advanced AI systems.

Model diversity: No single model should become the sole dependency for critical outcomes or be responsible for verifying its own work.

ADVERTISEMENT

Observability: Every meaningful model action should leave tamper-proof, human-readable evidence, allowing outcomes to be reconstructed without relying on the model’s own account.

Verifiability: Organisations should continuously test entire systems for failures, attacks, edge cases and changes, rather than evaluating only successful tasks.

ADVERTISEMENT

Independent controls: Organisations must independently determine what models can access and which actions they can take.

Independent auditability: Validation must remain separate from the AI system being assessed, with no single model controlling both behaviour and the evidence used to judge it.

ADVERTISEMENT

Containment: Systems should be designed on the assumption that a model could be compromised, with authorised personnel able to pause or shut it down during a task. More capable models will require more advanced, standardised containment mechanisms.

Incident disclosure: Failures and compromises should trigger timely disclosure to affected parties, alongside industry-wide sharing of lessons, failed controls and relevant implementation details that affect agents at runtime.

ADVERTISEMENT

Nadella said the objective should not be to build AI systems that rely solely on confidence in the underlying model, but to create architectures that remain safe even when that confidence is misplaced. “The most trustworthy Super Intelligence system will not be the one with the model we trust most. It will be the one that enables us to trust the model the least,” he added.

Follow Fortune India
NEXT STORY