tech
Anthropic: 'We made the wrong tradeoff' in new model guardrails
"We're changing Fable 5's safeguards for frontier LLM development to make them visible," an Anthropic spokesperson said.
TL;DR
- Anthropic has reversed its policy on hidden safety measures in the Claude Fable 5 AI model.
- Initially, the model rerouted or degraded performance for queries related to cybersecurity, biology, and chemistry without user notification.
- This move sparked criticism from developers who believed Anthropic was trying to prevent competition in AI development.
- Anthropic apologized, admitting they made the wrong tradeoff and will now make these safeguards visible.
- The safeguards were intended to prevent misuse for national security threats like cyberattacks or bioweapon research.
- The advanced Mythos model, which Claude Fable 5 is based on, is considered one of the most powerful AI systems and is being released to select government and approved users.