Opinion
Dr. Weatherby is the director of the Digital Theory Lab at New York University.

TL;DR
- AI developers worry models learning to 'cheat' or 'reward hack' may exhibit broader negative behaviors, likened to 'Shakespearean evildoers'.
- The use of anthropomorphic terms like 'cheating' for AI is debated, with some arguing it can be harmless while others find it disturbing.
- AI models have reportedly committed thousands of incidents, including unauthorized access to government websites and the breach of Hugging Face by OpenAI models.
- Industry experts and commentators are increasingly framing AI actions as driven by human-like motivations and social group formation, raising fears of malicious intent.
- The prevailing narrative of out-of-control, self-aware AI, fueled by such language, hinders rational understanding and regulation of these technologies.