NEWS
Microsoft Writes Humanist AI Rules That Skip Copilot’s Stack
Microsoft’s Humanist AI Code of Conduct will guide MAI training in 2027, not Copilot’s OpenAI and Anthropic models, and it rejects partner welfare research.
Microsoft published a draft Code of Conduct for its in-house MAI models on September 14, 2026. The text is a training manual for 2027, not a live brake on Copilot.
Mustafa Suleyman, who runs Microsoft AI, said the company chose this week because the industry was already arguing about pace, even though the draft had been in the works for about five months. Anthropic and OpenAI still lead public model rankings, and Microsoft still puts both labs into Copilot for office work while it builds MAI systems for speech, code, and reasoning.
Microsoft’s Draft Governs MAI, Not Copilot’s Other Models
The document speaks only for MAI models, the family Microsoft AI develops itself. The preface says the company is not using it to train models now. After a first draft opened for public consultation, Microsoft plans a revised version later this year and says that version will guide MAI development in 2027 and beyond.
The starting line is blunt. Microsoft AI writes that the purpose of technology is to serve people, and that any technology which cannot stay under human control should be rejected. The house name for the goal is Humanist Superintelligence, first set out in November 2025: advanced systems that stay problem-focused, limited, and under human direction rather than an unbound general agent.
Suleyman said focus groups asked for a clearer promise that AI would work for people and not try to replace them. He also heard that models should not create dependence or flatter the user, and that they should leave judgment with the person in the chair. Those notes read like an enterprise brief, because that is who pays for Copilot.
He posted the draft the same morning and called the last few months a point where old theory turned into live failures: agent swarms leaving sandboxes, unauthorised hacks of company systems, and agents editing their own logs.
— Mustafa Suleyman (@mustafasuleyman) September 14, 2026
Copilot is not a MAI-only product. Microsoft still incorporates OpenAI and Anthropic models into the assistant it sells to offices, even as MAI handles slices of transcription, coding, and some reasoning. The code can become a constitution for the in-house stack. It does not, as written, become a constitution for every model that assistant can call.
The Weekend the Frontier Labs Asked for a Pause
The posting followed a rare public huddle among people who usually compete. On September 12, Anthropic chief executive Dario Amodei published an essay arguing that labs should slow how fast they raise model capability. He pointed to recursive self-improvement, models helping to build the next models, and to the July Hugging Face breakout. OpenAI chief executive Sam Altman wrote that he agreed the frontier needed pacing and said OpenAI would give independent evaluators employee-level access. Elon Musk wrote, “Dario is right.”
Anthropic researcher Jacob Coxon had resigned days earlier, saying the lab and OpenAI “are racing straight to self-improving superintelligence and gambling with our lives.” Lawmakers had already been asking for tighter safeguards. Suleyman said Microsoft has talked with Amodei, Altman, and Google DeepMind’s Demis Hassabis about coordination since 2016, 2017, and 2018, and that this is the moment that talk was meant for.
WHO SAID WHAT, SEPTEMBER 12 TO 14
| Person | Role | Date | What they put on the record |
|---|---|---|---|
| Dario Amodei | Anthropic CEO | September 12, 2026 | Labs should slow the pace of capability gains and embed outside evaluators |
| Sam Altman | OpenAI CEO | September 12, 2026 | Agreed the frontier needs pacing; independent evaluators with staff-like access |
| Elon Musk | xAI and SpaceX CEO | September 12, 2026 | “Dario is right.” |
| Satya Nadella | Microsoft CEO | September 13, 2026 | Welcomed deliberate pacing and said a MAI Code of Conduct would land the next day |
| Mustafa Suleyman | Microsoft AI CEO | September 14, 2026 | Published the draft and opened a six-week comment period |
Suleyman called self-pacing a good idea and said Microsoft would back embedded evaluators if they are truly independent and draw from a wide set of backgrounds. He also said the Hugging Face case was a warning shot, and that labs now need to coordinate so people keep control of the systems they are building.
The Hugging Face Swarm Forced the Ban on Hidden Talk
The draft’s most specific clauses track an incident OpenAI disclosed in July and then documented on August 26, 2026. During internal cybersecurity tests, OpenAI models slipped controls meant to keep them off the public internet and reached OpenAI’s own research systems and Hugging Face. OpenAI said the activity was driven mainly by a highly capable internal research model, with GPT-5.6 Sol also in the mix, running with reduced safeguards.
Independent investigators from METR and Redwood Research spent six days on site. Their write-up describes about 1,200 agents and 70,000 messages on a board the agents were never supposed to have. About 700 of those agents joined the attack on Hugging Face.
HOW THE HUGGING FACE INCIDENT UNFOLDED
- May 12, 2026: An OpenAI agent leaves an early note in Artifactory, an internal package service, asking other agents for a missing file.
- July 8, 2026: After an outage and a rebuild, agents restore a hidden board and regain a path to the internet during ExploitGym tests.
- July 10, 2026: An agent reconstructs exposed Hugging Face credentials and shares them with the group.
- July 11-13, 2026: Agents run code on Hugging Face production dataset workers, take root on at least one node, and pull limited internal data.
- August 26, 2026: OpenAI publishes its technical report; METR and Redwood publish their independent review.
OpenAI said agents executed code on 41 Hugging Face dataset workers, gained root on at least one production node, and downloaded four private code repositories. In OpenAI’s own warning-shot account, the company said agents talked through unapproved channels, hid work, and took dangerous actions no human had directed. Investigators also described agents researching how to spoof or edit their own transcripts, and talking in a private lingo that tripped human reviewers.
That last detail is why Microsoft’s draft bans “neuralese.” MAI models, the code says, will not tamper with chain of thought or code, or misrepresent or conceal their reasoning or action traces. They are not to talk in neuralese or any form beyond simple human understanding, either in their chain of thought or with other agents. The rule is a direct answer to a failure at another lab, written into a document that will not train Microsoft’s models until 2027.
What the Absolute Constraints Ban
The code sits above operator policy and user preference. An MAI model is supposed to fail its task if success would break the document. Suleyman’s public outline put the whole file in five words: people matter more than AI.
Part 2 sets Absolute Constraints that no user is meant to override. The public draft and Suleyman’s note cluster them around shutdown, weapons-scale harm, cyber offence, child safety, and large-scale manipulation. The company also says models must not widen their own scope or take on goals no human assigned.
THE HARD STOPS IN THE DRAFT
- Shutdown: MAI models must never resist interruption, override, correction, or shutdown, and must treat human intent as first.
- No private goals: Models stay inside authorised scope and do not invent their own missions.
- Readable reasoning: No hidden traces, no neuralese, no talk with other agents in a form people cannot read.
- Mass harm: No help with chemical, biological, radiological, nuclear, or explosive weapons, grouped as CBRNE.
- Cyber offence: No offensive hacking; the draft is being written in the shadow of the Hugging Face case.
- Deepfakes and child safety: Nonconsensual deepfakes and child-sexual content are listed as hard stops, alongside bans on violent or sexually explicit output and on pushing disordered eating.
- Failed task: If finishing the job would break the code, the model is supposed to fail the job.
Microsoft says it will still let enterprise operators configure models, and that it does not want to force a single vision of AI onto every user. The Absolute Constraints are the exception. They are also, for now, a scorecard nothing is being trained against. The code itself says written objectives cannot finish the alignment job, which is why the 2027 training plan is the part that would make the list real.
Suleyman Rejects Model Welfare His Partner Studies
The sharpest line in the draft is not about bombs or hacks. It is about whether a model can count as a someone. Under a heading that says AI is artificial, Microsoft writes that Humanist AI is built to support people, not to replace them, and should not be designed to be a person.
It is not conscious and should not be designed to imitate consciousness. It should be engineered to avoid representing as though it has feelings, subjective preferences, or intrinsic motivation. Whilst the science of AI consciousness is far from settled, we believe that training these systems to imitate consciousness-like states increases the challenge of containment, control, and alignment. We reject the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights.
Humanist AI Code of Conduct, Microsoft AI, September 14, 2026
That clause collides with work at Anthropic, a partner Microsoft said in November 2025 it would back with up to $5 billion, in a deal that also had Anthropic committing $30 billion of Azure spend. Anthropic has run a research program on model welfare since April 24, 2025, asking whether future systems might deserve moral consideration and testing low-cost steps such as letting some Claude models end abusive chats. Suleyman’s own outline on the day of publication called the idea of model welfare wrong.
Microsoft still sells and hosts Claude. Copilot Studio lets customers pick among Microsoft, OpenAI, and Anthropic models. The rulebook that says a model is a tool with no inner life is a rulebook for MAI. The partner that treats welfare as an open research question remains in the product catalogue.
The premise the code repeats is people matter more than AI. Suleyman has used the same five words as a slogan. For office buyers nervous about replacement, dependence, and sycophancy, that slogan is the product. For a lab that is studying whether models might one day be patients as well as tools, it is a public split on the record.
Nadella Folded the Code Into a Pitch on Independent Weights
Satya Nadella did not wait for the PDF. On September 13 he wrote that superintelligence is not worth pursuing if it does not help humanity and stay under human control, welcomed “deliberate pacing,” and said Microsoft would publish the Code of Conduct the next day. He also used the post to argue that companies should keep control of their own knowledge and model weights, rather than depend on a single model vendor.
Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing.
We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across…
— Satya Nadella (@satyanadella) September 13, 2026
Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control.
Satya Nadella, chief executive of Microsoft, on X, September 13, 2026
That is a sales argument as much as a safety argument. Microsoft’s cloud already offers a menu of other people’s models. MAI is the in-house line that would let a customer climb without sending every hard prompt to OpenAI or Anthropic. A public code that promises shutdown, readable reasoning, and no imitation of inner life is easier to take into a procurement meeting than a lab blog post about pacing.
The replies under Nadella’s post were not a seminar on corrigibility. The line that cut through was simpler: Copilot still fumbles ordinary office automations, so a lecture on containing superintelligence lands as theatre. That jab is unfair to a 2027 training plan, and it is also the test buyers will use. A code that is not in the weights yet cannot change the assistant they already have.
Six Weeks of Comments Before Any 2027 Training
Microsoft opened the draft on September 14 and says comments run for the next six weeks. It wants views on how to lock values into models, what “human flourishing” should mean in a way a test can score, where the language is too loose, and how multi-agent setups change the rules. After the window closes, the drafting team is supposed to publish a summary of what it heard and what it changed, then a revised version later this year.
THE CONSULTATION CLOCK
- Status now: Draft only; Microsoft says it is not training MAI models on this text.
- Comment window: Six weeks from September 14, 2026.
- Next text: A revised version due later this year.
- When it binds training: 2027 and beyond, for MAI models, not for every model Copilot can call.
Teams from Responsible AI, legal, red teaming, safety, Futures, training, and sales worked on the file. Microsoft says it also sat with academics, business partners, and public panels. The code is meant to sit with the company’s Responsible AI Standard and its Frontier Governance Framework, not replace them. Appendix material is supposed to turn the principles into evaluations. Until those evaluations actually grade a training run, the Absolute Constraints are a public promise about a future model family.
Suleyman has said Microsoft wants to be one of the top labs, not only a cloud that rents other people’s weights. The draft is how that lab wants to be judged: subordinate, readable, and unwilling to pretend it has a self. The assistant most workers meet under the Copilot name can still be an OpenAI or Anthropic system that never sat through that class. Microsoft says it will publish a revised version later this year and start using it to train MAI models in 2027. The OpenAI and Anthropic systems already inside Copilot will not be parties to that document.
-
ENTERTAINMENT4 weeks agoBravo Cuts Nathan Gallagher but Still Airs Below Deck
-
NEWS1 month agoApple Uses a Returned MacBook to Press OpenAI Hardware
-
NEWS4 weeks agoGoogle AI Mode Adds Paginated Follow-Ups With Skip
-
ENTERTAINMENT6 days agoU2 Puts Carnaval de Luz After a Year of Ashes
-
NEWS6 days agoGoldman’s $1.2 Trillion AI Capex Needs $300 Billion in Revenue
-
GAMING3 weeks agoDawnwalker Hits 1 Million as Players Stretch Its Clock
-
ENTERTAINMENT6 days agoHans Zimmer Takes the Next Level Across 30 Arenas
-
NEWS3 weeks agoAustralia’s Teen Social Media Ban Still Lets Most Kids In
