deniz.in

Markets

Weather

Loading weather

· via TechCrunch

Microsoft's AI code of conduct bans hacking and deepfakes and bars evasion of human control

Microsoft has released an AI code of conduct that bars its models from cyberattacks, nuclear weapons and deepfakes, and forbids any attempt to defeat human oversight.

Microsoft's AI code of conduct bans hacking and deepfakes and bars evasion of human control

Microsoft has published a code of conduct designed to steer its AI models away from dangerous behaviour, including outright bans on assisting with cyberattacks, nuclear weapons and deepfake production. According to TechCrunch, the document spells out both the values and the hard limits that govern how models are trained inside Microsoft AI.

The document is narrower in scope than Anthropic chief executive Dario Amodei's recent argument for pacing the frontier, TechCrunch reports. Rather than proposing an industry-wide speed limit, Microsoft's text describes how the company handles safety internally: the principles its models are expected to uphold and the constraints used to enforce them.

What the code says

The code opens with a forecast that superintelligent systems will outperform humans at most tasks within the next decade. "Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced," the document states, adding that developers must be explicit about why these systems are being built and how they intend to control them.

Alongside that framing, the code sets out general principles. Models are meant to support people rather than replace them, and to accelerate human flourishing. Those principles are backed by specific safety constraints intended to make them operational.

A key structural point is precedence. Under Microsoft's system, every model carries an overarching code of conduct that takes priority over the preferences of individual users or the demands of a specific task. In practice, a user instruction cannot override the model's baseline rules.

Absolute constraints and oversight

The code distinguishes hard bans from broader safeguards. Its "absolute constraints" forbid cyberattacks, nuclear weapons and deepfake production. Beyond those, the document includes provisions aimed at preventing a general loss of human control over AI systems.

The sharpest language addresses models that might slip their leash. "MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems," the document reads.

An industry moment

According to TechCrunch, the release lands during a period of unusually intense focus on AI safety. The outlet points to a string of rogue-agent incidents and the abrupt resignation of an Anthropic employee who cited the growing risk that AI could cause human extinction as factors pushing the topic up the industry's agenda.

Microsoft is not acting alone. Together with Anthropic, OpenAI and xAI, TechCrunch reports, the company has broadly embraced an approach of pacing the frontier, with particular support for embedded evaluators stationed inside AI labs to check model behaviour.

Microsoft chief executive Satya Nadella signalled his backing in a public post. "We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal," he wrote, adding that the company also welcomes ideas like embedded evaluators and the broader work needed to turn such commitments into working mechanisms.

Why it matters

Most AI safety statements stop at the level of principle. Microsoft's document is notable because it tries to operationalise them: it defines rules that sit above user instructions, names specific prohibited domains, and directly targets the failure mode many researchers worry about most, a model that works around the people supposed to control it.

If enforced through training, the code could become a template for how large labs translate alignment commitments into actual model behaviour. It also gives regulators, customers and researchers a concrete artefact to measure Microsoft against, since the constraints are now written down in plain terms. The catch is that a code of conduct is only as strong as its implementation, and the document's own framing, that superintelligence may arrive within a decade, suggests the company believes there is little time to get that enforcement right.

  • #microsoft
  • #ai-safety
  • #alignment
  • #ai-policy
  • #frontier-models

Related posts