tech

Anthropic's 'safe' Mythos-class model won't answer questions about cancer

The broad safeguards built into Anthropic's Claude Fable 5, a Mythos-class AI model, blocks some mundane requests on cybersecurity and biology.

Anthropic's 'safe' Mythos-class model won't answer questions about cancer

TL;DR

  • Anthropic's Claude Fable 5, a Mythos-class AI model, includes broad safeguards that can flag and block requests on cybersecurity and biology.
  • These safeguards were implemented to allow the powerful model's public release while preventing potential misuse in sensitive areas like biological research.
  • When safeguards are triggered, Fable 5 may refuse to answer or switch to a less capable model, Opus 4.8.
  • Anthropic acknowledges that safe content may be flagged and is working to improve the safeguards to reduce false positives.
  • The company intends to eventually release Mythos-class models without these stringent safeguards for the broader scientific community to accelerate research.