Crying Safety Got It Banned
An unprecedented situation is unfolding in which the company that has spoken most openly about safety in the AI industry has been tripped up by its own honesty.
On June 12, 2026, Anthropic revealed in an official blog post that it had received an order from U.S. government authorities to suspend access to its latest and most powerful AI models, "Fable 5" and "Mythos 5." The reason cited was "limited jailbreak potential"—the discovery of methods capable of circumventing the models' safety controls.
What makes this significant is that Anthropic itself was the first to report the vulnerability. The company has long championed its "Responsible Scaling Policy (RSP)," taking a proactive stance of internally evaluating and actively disclosing risks associated with its models. This time, information was provided to authorities as part of that same process—but the government used that report as the basis for ordering a halt to commercial deployment.
Anthropic's Rebuttal — "Not a Reason to Recall a Model Used by Hundreds of Millions"
In its statement, Anthropic clearly pushed back against the government's decision. The company argued that the measure was an overreaction, stating that it does not believe "the discovery of a limited jailbreak constitutes grounds for recalling a commercial model deployed to hundreds of millions of users."
What Anthropic takes issue with is the rigidity of the regulation. Some vulnerability exists in every model. Providing complete proof of safety is nearly impossible at this stage. And yet, if voluntary disclosure leads immediately to suspension of operations, companies will gradually develop an incentive to not report honestly—Anthropic is sounding the alarm about that paradox for the entire AI industry.
The Developer Community Is Shaken
The news sparked extensive debate on Hacker News as well. Among developers, there have been reports of cases where developers who had experimentally launched Fable-based games—one called "Shepherd's Dog"—were affected by the access suspension, and the real-world impact on commercial use is spreading.
The independent blog "12 Grams of Carbon," in an analysis piece titled "A Large Shadow Hangs Over This Fable Issue," points out that the incident presents not merely a regulatory problem, but a fundamental question: "How much transparency must an AI company demonstrate before it is considered safe?"
Does Disclosing Safety Risks Become a "Punishment"?
This issue cuts to the heart of a fundamental dilemma in AI regulatory design. What governments and society demand of AI companies is transparency and the assurance of safety. But if complying with that transparency carries the risk of having a model shut down, how will companies behave?
In future AI regulatory discussions, establishing the principle that "self-reporting of vulnerabilities should not result in penalties" will be an urgent priority. Anthropic's case looks set to become a catalyst that accelerates that conversation.
Summary
The safety concerns that Anthropic itself reported have come back to bite it in the form of suspended access to Fable and Mythos. This is not merely Anthropic's problem—it is an issue concerning the design of transparency incentives for the entire AI industry. If the structure in which "the company that speaks honestly loses out" is left unaddressed, the industry as a whole may become reluctant to disclose information. A more sophisticated framework for handling vulnerabilities between regulators and AI companies is urgently needed.