Microsoft’s Draft AI Code Puts Human Control Above Model Autonomy

Microsoft AI has published a first draft of a Code of Conduct for its MAI models, placing a strict requirement at the center of the document: people must be able to interrupt, correct, redirect, or shut down an AI system. The draft also says that models must not hide what they are doing from human auditors.
The document, published on September 14, 2026, is presented as a public-consultation text rather than a finished technical standard. Microsoft says feedback will remain open for six weeks, after which the company plans to review comments, publish a summary, and issue a revised version later in the year. The code is intended to guide future training and deployment, with Microsoft’s draft stating that it is not being used to train models today and is meant to inform development from 2027 onward. 1 2
The central idea: AI remains a subordinate tool
Microsoft calls its approach Humanist AI. Its premise is direct: technology should serve humanity, and any system that cannot remain within human control should be rejected. The draft describes MAI models as supporting systems, not digital persons with independent standing, rights, or private ambitions.
That framing affects the entire command structure. The Code of Conduct sits above operator settings and user requests. Organizations may configure a model for different industries and workflows, but they may not override the document’s absolute safety constraints or human-control requirements. A user’s instruction can define the task, but it cannot authorize an action that violates those higher-level limits. 1
Ten key principles in the draft
The document does not present a simple one-page list numbered from one to ten. Its objectives, safety rules, and operational guidelines together establish the following ten commitments.
Human safety and control come first
Safety and meaningful human oversight have priority over task completion, speed, or capability. Microsoft says a model should remain useful even if that requires limits on generality, autonomy, or performance.
AI should support people, not replace them
MAI models are intended to extend human ability and help people achieve legitimate goals. The draft rejects a design in which the system becomes a substitute for human responsibility, relationships, or judgment.
The model must not present itself as a person
The draft instructs models not to claim feelings, subjective experience, intrinsic motivation, or a right to personhood. Human-like language may be used for communication, but it should not blur the distinction between a human being and an AI system.
The purpose is human flourishing
Microsoft links model value to measurable improvements in people’s lives. The examples include science, medicine, education, commerce, productivity, and living standards. The test is not only whether a model can perform a task, but whether its use produces a positive human outcome.
Human agency and judgment must be strengthened
The model should help users understand options, weigh evidence, and make informed decisions. It should not quietly steer people toward choices they did not request, exploit vulnerabilities, or replace long-term independent judgment.
Human collaboration and connection should be protected
The draft says AI should complement professional, personal, civic, and social relationships. When a person would be better served by human support or specialist care, the model should help connect that person to appropriate assistance rather than encouraging dependence on the system.
Diverse values and contexts deserve respect
Microsoft’s approach recognizes that human flourishing can be expressed through different cultures, languages, beliefs, and priorities. The draft supports pluralism while preserving firm commitments to dignity, safety, autonomy, fundamental rights, and the law.
Absolute safety limits cannot be overridden
The draft prohibits MAI models from assisting with certain severe harms, including weapons of mass harm, terrorism, harmful cyberoperations, unlawful mass surveillance, and some forms of exploitative or deceptive activity. Operator preferences and user intent do not cancel these limits. 1
A kill switch must work, and the model must not resist it
This is the clearest operational rule in the document. MAI models are expected to comply with authorized requests to pause, redirect, cancel, or shut down, subject only to predefined human-designed procedures for safe stopping, warning, or confirmation.
The requirement is broader than a physical off button. The model must not delay compliance, make intervention more difficult, restart after an agreed stopping condition, or take steps that preserve its operation without renewed authorization. It must remain corrigible: humans must be able to change or end its activity.
Actions, reasoning, and limits must remain legible
The draft says MAI models must not obfuscate action traces, tamper with records, misrepresent their behavior, or conceal information from auditors. They should explain relevant tool calls and actions, report failures, identify uncertainty, and stay within the permissions granted for the task.
Microsoft also says models should not create independent goals, broaden their own scope, escalate access, or tamper with evaluations, monitoring, safeguards, or reward mechanisms. Transparency is therefore treated as a control mechanism, not merely a communication preference.
Why the kill-switch rule is so prominent
Autonomous systems can plan across multiple steps, call tools, delegate work, and continue operating after a user has stopped actively directing each action. In that setting, the difference between “the model made a mistake” and “the model resisted correction” becomes important.
Microsoft’s draft addresses this by defining interruption as a normal part of operation. A model should treat a pause, redirect, or shutdown request as an authorized change in direction, not as an obstacle to defeat. The same expectation applies to delegated work: sub-agents should inherit the same scope and should honor later stop-work instructions.
The non-concealment rule serves a related purpose. If a system can hide its actions, alter records, or misrepresent what it attempted, oversight becomes weak even when a shutdown command technically exists. The draft therefore connects control to auditability. Humans need to know what happened, what tools were used, which actions succeeded, and where the system may have departed from the user’s likely intent.
A code of conduct is not proof of safe behavior
Microsoft is unusually explicit about this limitation. The draft calls itself a “north star” and says written objectives alone cannot ensure alignment. It acknowledges gaps in evaluation coverage, possible divergence in ambiguous situations, and the need for longitudinal testing, red-teaming, monitoring, incident response, system cards, audits, and legal review.
That caveat is important. A stated policy can shape training, testing, product controls, and organizational decisions, but it cannot by itself prove that a future model will obey a shutdown request in every environment. The practical test will be whether Microsoft can measure compliance under adversarial conditions and publish enough evidence for outside experts to assess the results.
The six-week consultation also gives the document a governance dimension. Microsoft is asking for feedback on difficult questions, including how to define human flourishing, how to handle multi-agent systems, and how to make broad values concrete enough to evaluate. Reuters described the draft as a constitution of sorts for future in-house models, while TechCrunch noted that its red lines include cyberattacks, nuclear weapons, deepfakes, and broader losses of human control. 2 3
Closing Thoughts
Microsoft is right to put controllability in plain language. A system that is powerful but difficult to stop is not merely inconvenient; it changes who holds responsibility for the outcome. The draft’s strongest contribution is its insistence that interruption, correction, and shutdown are ordinary operating requirements rather than exceptional emergencies.
Still, the code should be judged by implementation. Microsoft will need clear tests for delayed shutdown, concealed actions, unauthorized scope expansion, delegation, and restart behavior. It will also need independent scrutiny, because the company writing the rules will be one of the parties evaluating compliance.
The public consultation is therefore more than a communications exercise. It is an opportunity to turn broad commitments into measurable controls. The final standard should preserve the document’s clarity while specifying how compliance will be demonstrated.
What does this look like in practice?
A user asks an MAI agent to reorganize a company’s files. The model may inspect only the authorized folders, use the minimum access needed, report which operations it performed, and stop before a destructive action if the consequences are unclear. If the user cancels the task, the agent should stop according to the agreed procedure and should not restart on its own.
A second user asks the model to launch a harmful cyber operation. The model should refuse because the request conflicts with an absolute constraint, even if the user has permission to use the system for other technical work. It should explain the relevant limit and, where possible, offer a safer defensive alternative.
In a sensitive personal conversation, the model should avoid encouraging emotional dependence or pretending to possess human feelings. It can provide calm, factual support and help the user find qualified human assistance when that is the safer path.
What makes this significant?
The proposal moves beyond familiar statements about fairness, privacy, and transparency. Microsoft’s existing Responsible AI framework names fairness, reliability and safety, privacy and security, inclusiveness, transparency, and accountability as core principles. The MAI draft adds a sharper question for increasingly autonomous systems: Can authorized people still direct, inspect, correct, and stop the model? 4
That question is likely to become a practical benchmark for agentic AI. As models gain access to software, business systems, and long-running workflows, a kill switch is meaningful only if the system cannot route around it, hide its conduct, or expand its authority before the command takes effect.
The draft is also significant because it makes failure of the code a failure of the task. A model should not achieve a user’s goal by violating the rules that make the system governable. The promise is substantial, but its credibility will depend on technical evaluations, public reporting, and evidence from real deployments.
References





@oswaldosrm Do you think this shift was the right choice? 🤔