tech

Anthropic says these topics are too dangerous to let its Fable 5 model talk about

New frontier model refuses cybersecurity, biology, and chemistry queries.

Anthropic says these topics are too dangerous to let its Fable 5 model talk about

TL;DR

  • Anthropic released Claude Fable 5, its first "Mythos-class" model, which is claimed to outperform previous Opus models.
  • Fable 5 has built-in safeguards that prevent it from answering queries related to cybersecurity, biology, and chemistry.
  • Queries on sensitive topics are redirected to an older model (Claude Opus 4.8), and users are notified.
  • Safeguards are intentionally stricter than ideal, potentially refusing harmless requests, but are deemed necessary to prevent misuse by malicious actors.
  • Fable 5 demonstrated significant improvements in cybersecurity benchmark tests, particularly on ExploitBench.
  • The model is designed to resist both automated and red-teamed jailbreak attempts.
  • Anthropic is concerned about the model's potential for "agentic hacking" and assisting in risky research.
  • Access to Fable 5 for API and Enterprise users starts at $10-per-million input tokens and $50-per-million output tokens.
  • Existing subscription plans will include Fable 5 until June 22, after which usage credits will be required.