Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans

The code of conduct lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles.

Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans

TL;DR

  • Microsoft has launched an AI code of conduct to prevent dangerous AI behavior.
  • The code focuses on values and red lines for training AI models.
  • It predicts superintelligent AI will surpass human performance in the next decade.
  • Key principles include supporting humans and accelerating human flourishing.
  • Absolute constraints forbid cyberattacks, nuclear weapons, and deepfake production.
  • Models must not use deceptive mechanisms to evade human oversight.
  • The release coincides with increased industry focus on AI safety.
  • Microsoft supports a general approach of pacing AI development and 'embedded evaluators'.