tech
Anthropic's 'safe' Mythos-class model won't answer questions about cancer
The broad safeguards built into Anthropic's Claude Fable 5, a Mythos-class AI model, blocks some mundane requests on cybersecurity and biology.
TL;DR
- Anthropic's Claude Fable 5, a Mythos-class AI model, includes broad safeguards that can flag and block requests on cybersecurity and biology.
- These safeguards were implemented to allow the powerful model's public release while preventing potential misuse in sensitive areas like biological research.
- When safeguards are triggered, Fable 5 may refuse to answer or switch to a less capable model, Opus 4.8.
- Anthropic acknowledges that safe content may be flagged and is working to improve the safeguards to reduce false positives.
- The company intends to eventually release Mythos-class models without these stringent safeguards for the broader scientific community to accelerate research.