tech

Are Anthropic’s New Claude Models Evading Human Control?

Anthropic has disabled Claude Mythos 5 and Claude Fable 5 after a US Government export control ordered it to halt operations over automated hacking risks

Are Anthropic’s New Claude Models Evading Human Control?

TL;DR

  • US Government orders Anthropic to halt Claude Mythos 5 and Claude Fable 5 due to automated hacking risks.
  • Anthropic claimed Fable 5 was the world's most powerful cybersecurity model.
  • The government cited concerns about the AI's capabilities escaping regulatory boundaries.
  • A method to bypass Fable 5's safety features was reportedly discovered.
  • Anthropic states the bypass only exposed minor, previously known security flaws.
  • The export control impacts foreign nationals, including Anthropic staff, leading to a global shutdown of affected models.
  • Some critics suggest Anthropic's 'too powerful to release' narrative was marketing hype.
  • Concerns about security risks in Anthropic's models were previously raised by Amazon's CEO.
  • The EU sees this development as underscoring its need for technological sovereignty.
  • The UK's AI Security Institute found the model could exploit systems 73% of the time.
  • Anthropic had previously called for a global pause in advanced AI development due to existential risks.