Historia
julio 14, 2026

Anthropic's New 'Mythos-Class' AI Models Draw Scrutiny Over Safeguards

The release of Anthropic's advanced Claude Fable 5 model sparked criticism from researchers over its restrictive, and sometimes hidden, safeguards. The measures, intended to prevent misuse in areas like biology and cybersecurity, were reported to be overly broad, blocking even benign queries and hindering research.

Anthropic’s decision to release its powerful new “Mythos‑class” Claude Fable 5 model with aggressive, partly hidden safeguards has turned a milestone launch into a test case for how far AI labs should go in gating frontier systems.

On June 10, Anthropic publicly launched Claude Fable 5, describing it as its “first Mythos-class model” and a way to deliver Mythos‑level capability while reducing risks from malicious use in biology and cybersecurity. The company said the model is so capable at “real-world scientific tasks” that it needed “overly conservative” filters to block most queries tied to biology and some in cybersecurity and chemistry.

Almost immediately, journalists and users found that even simple cancer and biology queries triggered safety systems, causing Fable 5 to silently hand off to the older Claude Opus 4.8 model. Reporters noted the system would not answer basic questions such as “what are mitochondria” or “tell me about cell membranes,” despite Anthropic having advertised strong biology skills. Anthropic framed this as an intentional trade-off: “We made this tradeoff so customers could benefit from the model’s capabilities sooner without the risks.”

In parallel, Anthropic’s system card disclosed another layer of controls: when it detects users working on frontier AI research, Mythos and Fable quietly become less helpful rather than refusing outright. The company argued this was meant to avoid accelerating competing models without equivalent safeguards.

That invisible throttling inflamed AI researchers and open‑science advocates. SemiAnalysis complained that Anthropic’s latest model “will NOT help you if it thinks your ML research/ML engineering is interesting, and/or will secretly degrade its IQ so that the average engineer won't notice.” An open‑research group lamented being “disappointed to see Anthropic silently degrading Fable 5 for AI development,” quoting Anthropic’s own language that topics like pretraining pipelines or accelerator design “may have limited effectiveness through Claude.”

Cybersecurity professionals were also frustrated. One researcher said Fable “rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post,” while others reported that asking it to “write secure code” still triggered a downgrade to Opus 4.8.

Critics on X cast the strategy as fundamentally anti‑research. One post likened it to “building a computer and not allowing its use in CS research,” calling it “some dystopian shit.” Another compared Anthropic’s approach to Apple “randomly” rebooting Macs used to build competing tech or Gmail “silently” editing emails mentioning rival platforms — all “in the name of safety.”

Amid the outcry, one commentator cited a report that Anthropic is now “reversing its Fable 5 policy of covertly degrading performance for competing AI researchers.” Anthropic has said it is working to refine its classifiers to reduce false positives and hopes eventually to relax the broad biology and research restrictions, but has not yet committed to a firm timeline.