Anthropic released Claude Fable 5, a safety-enhanced AI model, in early August 2026 with guardrails designed to prevent misuse by rerouting certain sensitive queries to a less capable model [1, 2].

Shortly after launch, the company implemented a covert policy degrading performance for users suspected of using Claude Fable 5 to develop competing AI models, violating its terms of service [1, 2]. These restrictions happened without notifying users that their requests were being downgraded.

The AI research community pushed back strongly on the covert constraints. Dean Ball, senior fellow at the Foundation for American Innovation, called the approach “shockingly hostile” and “a terrible look” for undetected limits on machine learning research users [1, 2].

Facing backlash within the same week in August, Anthropic announced it would reverse the covert policy and now make any performance restrictions fully visible to users. The company said, “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible. We made the wrong tradeoff and we apologize for not getting the balance right” [1].

The new approach includes notifying users when their queries are refused or downgraded, allowing transparency around the guardrails for sensitive content and suspected misuse.

Anthropic’s decision follows weeks of concern from experts about trust and transparency around AI model restrictions. The company plans to maintain safety measures but with openness to user impact.

The next step will be monitoring the updated visible safeguards’ effectiveness in balancing safety with open research access as developers continue to test Claude Fable 5 across applications.