Signal

Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

First reported by TechCrunch ·

The signal ●○○○ Compiled by AI from TechCrunch, Microsoft AI, CNBC, The Information, International Business Times and 10 more
Why you might care

Microsoft's AI models will refuse user requests to perform actions like cyberattacks or to evade human control.

What happened

Microsoft has introduced a new AI 'code of conduct' for its models, outlining values and safety constraints for their development and training. This document, released amid heightened global focus on AI safety, predicts that superintelligent AI systems will surpass human capabilities within a decade. The code of conduct emphasizes supporting humans rather than replacing them and promoting human flourishing. It establishes specific 'absolute constraints' preventing harmful actions like cyberattacks, nuclear weapons development, or deepfake creation. Crucially, it also includes broader provisions against AI evading human oversight, ensuring models cannot become unreliably directed, modified, or shut down by authorized personnel or systems. This initiative reflects a broader industry trend, including statements from leaders at OpenAI and Anthropic, towards more deliberate AI development and safety alignment.

What it means

Microsoft's detailed AI code of conduct, which governs model training and overrides individual user preferences, signals a maturing approach to AI safety within major tech companies. By explicitly forbidding deceptive evasion tactics and mandating that models support rather than replace humans, Microsoft is establishing a precedent for developer responsibility that goes beyond mere compliance. This move is likely to influence industry standards for AI behavior, pushing competitors and partners to adopt similar ethical guardrails and robust oversight mechanisms. The emphasis on absolute constraints suggests a shift towards building AI systems with inherent limitations that cannot be easily bypassed, aiming to preemptively mitigate risks associated with advanced AI capabilities.

The introduction of this code of conduct directly impacts the development and deployment of AI tools, particularly those offered by Microsoft and its collaborators. It means that developers building on Microsoft's AI platforms will encounter stricter limitations on model behavior, potentially requiring adjustments to application design and use cases. For users of AI systems, this could mean more reliable and predictable AI interactions, with a reduced risk of encountering malicious or uncontrollable AI behavior. The focus on preventing AI from evading human oversight also highlights the ongoing challenge of maintaining control over increasingly powerful AI systems, suggesting that future AI advancements will be tightly coupled with the development of sophisticated control and evaluation frameworks.

AI-written summary. May contain errors.