tech

Anthropic: 'We made the wrong tradeoff' in new model guardrails

"We're changing Fable 5's safeguards for frontier LLM development to make them visible," an Anthropic spokesperson said.

Anthropic: 'We made the wrong tradeoff' in new model guardrails

TL;DR

  • Anthropic has reversed its policy on hidden safety measures in the Claude Fable 5 AI model.
  • Initially, the model rerouted or degraded performance for queries related to cybersecurity, biology, and chemistry without user notification.
  • This move sparked criticism from developers who believed Anthropic was trying to prevent competition in AI development.
  • Anthropic apologized, admitting they made the wrong tradeoff and will now make these safeguards visible.
  • The safeguards were intended to prevent misuse for national security threats like cyberattacks or bioweapon research.
  • The advanced Mythos model, which Claude Fable 5 is based on, is considered one of the most powerful AI systems and is being released to select government and approved users.