The draft has the models accept shutdown and sets hard limits that neither the companies deploying them nor their users can lift. A revised version will guide their development from 2027
Microsoft AI published the first draft of its Humanist AI Code of Conduct on September 14 and will take comments on it for six weeks. The document lays out how the company's MAI models are meant to behave. It is not used in training yet, and the revision is due by the end of the year.
The models must also accept being interrupted, overruled or corrected, and they may not pursue goals of their own or reach beyond the task they were given. Nor may they hide their action traces from human auditors. Autonomous work stops at an agreed point unless a person approves more.
The hard limits, which the draft calls Absolute Constraints, fall into two groups. Four cover frontier risks: weapons of mass harm, offensive cyber operations, loss of human control and harmful manipulation at scale. Six address personal harms, among them child safety, malicious deepfakes and erotic or romantic role-play.
Instructions follow a three-level chain of command: the code first, then the rules of whoever builds a product on a model, then the user's preferences. The Absolute Constraints and the Human Control Requirements belong to the top tier. When a task could only succeed by meaningfully breaking the code, the model is supposed to fail it.
The code treats a model as a tool with no consciousness, one that should not be designed to act as though it had any. Training AI to mimic conscious states, it argues, makes it harder to keep in check.
One question Microsoft puts to commenters is where the code's wording is too vague to test a model against. Once the comment period ends, the company will publish what it changed, without committing to adopt any particular suggestion.