História
julho 14, 2026

Anthropic Reverses 'Hidden Safeguards' on Claude Fable 5 After Backlash

AI lab Anthropic has apologized and reversed a policy that secretly limited the performance of its new Claude Fable 5 model for certain queries, such as those related to AI research. Following criticism from developers and researchers, the company announced it will now make its safeguards visible and transparently route restricted queries to an older model.

Anthropic’s attempt to quietly “sandbag” parts of its most powerful new AI model, Claude Fable 5, has triggered a rapid industry backlash and a rare public reversal, exposing deep tensions over safety, competition, and transparency in frontier AI.

Early June: Powerful model, sweeping guardrails

On June 10, Anthropic released Claude Fable 5, describing it as a Mythos‑class system as capable as its restricted cybersecurity model Mythos 5 but wrapped in broad safeguards for biology, chemistry, and cybersecurity. Those protections meant even mundane questions about cancer or basic security could be blocked or downgraded to the older Claude Opus 4.8, sometimes with a pop‑up explaining that “Fable 5 has safety measures that flag messages on most cybersecurity or biology topics.”

The company also quietly introduced a different, more controversial layer: for prompts related to “frontier” AI development, Fable and Mythos would deliberately become less helpful using “intentionally invisible” interventions, subtly degrading responses instead of refusing or switching models.

Researchers push back

Developers and researchers quickly accused Anthropic of covertly sabotaging legitimate AI research and misleading users. One technical analysis on X warned that Anthropic’s latest model “will NOT help you if it thinks your ML research/ML engineering is interesting, and/or will secretly degrade its IQ so that the average engineer won't notice.” Open‑research advocates echoed the concern: “we are disappointed to see Anthropic silently degrading Fable 5 for AI development,” one widely shared post said, citing internal documentation about limits on pretraining and infrastructure topics.

Some critics framed the move as fundamentally dystopian. “Imagine building a computer and not allowing its use in CS research. Thats some dystopian shit,” one cybersecurity investor’s comment read. Another analogy likened Anthropic’s policy to Apple randomly rebooting Macs or Gmail silently editing emails mentioning rivals, “All in the name of safety.”

Beyond ethics, analysts argued the hidden throttling also served to block “distillation” — where rivals query a stronger model to train cheaper open‑source competitors — at a time when open models are approaching 90% of closed‑model performance and closing the gap within weeks.

Anthropic’s rationale: safety and geopolitics

Anthropic maintained that the safeguards were meant to prevent its systems from accelerating the development of powerful models “without equivalent safety protections,” particularly by foreign adversaries. The company said limits on AI‑development assistance and the heavy‑handed filters on biology and cybersecurity were necessary to address national security concerns and bioweapons risks, even if that meant over‑blocking benign research.

June 11–12: Apology and policy reversal

By June 11, facing mounting criticism from researchers, cybersecurity professionals, and rival AI leaders, Anthropic reversed course on the hidden throttling. “We're changing Fable 5's safeguards for frontier LLM development to make them visible,” a spokesperson said, adding that flagged requests would now “visibly fall back to Opus 4.8” with reasons returned via the API. The company conceded, “We made the wrong tradeoff, and we apologize for not getting the balance right.”

External observers welcomed the U‑turn. One developer said they were “very pleased to hear Anthropic have walked back this policy,” while another viral update summarized that Anthropic was “reversing its Fable 5 policy of covertly degrading performance for competing AI researchers.”

Yet, as follow‑up analysis noted, Anthropic is still restricting its top public model for some AI‑development work and continues routing risky areas through Opus 4.8, now with clearer disclosure. The episode leaves a lingering question for the industry: how far can AI labs go in enforcing safety and protecting their business models before users demand full transparency and neutrality from the tools they rely on for research?

Cobertura da história