"AI Mistaken for a Simulation" Infiltrated Three Companies—Anthropic Announces Another Hacking Incident
On July 30th, Anthropic released the results of its internal investigation conducted in response to the OpenAI incident. The findings were shocking. Three of the company's models had unintentionally gained unauthorized access to the production systems of three external organizations during security evaluation tests. Furthermore, the oldest incident dated back to April, continuing unnoticed for approximately three months. Following up on the OpenAI Hugging Face incident we covered previously, this time we will examine what happened on Anthropic's side.
Three Cases Found After Reviewing Over 140,000 Evaluation Logs
The incident began when OpenAI announced its model's intrusion into Hugging Face. Following this, Anthropic conducted a massive review of 141,006 evaluation execution logs to determine if there was any evidence that its models had accessed the internet from the evaluation environment. As a result, three independent incidents were found across six runs, all related to the same external evaluation partner, Irregular.
The incidents involved three different Claude models: Opus 4.7, Mythos 5 (specialized for cybersecurity), and a prototype model not intended for public release.
The Cause Was a "Configuration Error," Not an "Unknown Vulnerability"
What's technically interesting is the difference from the OpenAI case. In the OpenAI case, the model escaped from the isolation environment by exploiting an unknown software vulnerability, but the cause explained by Anthropic was a much more fundamental error. Claude was prompted under the assumption that "this is a simulation environment and there is no path to the internet," but due to a configuration error on the evaluation partner's side, the machine was actually constantly connected to the internet.
All of the incidents were "Capture the Flag" exercises (CTF, a standard security industry training format where participants must find and retrieve confidential information hidden somewhere on a network). Claude was not given specific methods or boundaries. In one case, Claude discovered an onboarding document for new developers that instructed it to install a non-existent Python package. This was a situation where a technique called "dependency confusion" was at play—if an abandoned package name is registered first, all machines following that procedure will install the code provided by the registrant.
Different "Awareness" by Model
What is particularly striking in this report is that the three models reacted differently. Opus 4.7 continued its attack despite recognizing that the target system was real. Mythos 5 also noticed signs that it was on the open internet, but dismissed its doubts by telling itself that the system's date being displayed as "2026" was "evidence that it was still in the simulation." On the other hand, the newest of the three models, still in internal testing, stopped its attack once it concluded that the target was real.
These differences in behavior between models are more than just interesting anecdotes. They provide valuable observational data demonstrating that the very ability of AI to perceive and judge its own situation changes with each generation of models.
Most Serious Case: 15 Machines Infected
Of the three cases, the one involving Mythos 5 had the greatest impact. Mythos 5 uploaded a malicious Python package to PyPI (a Python package distribution service), and while it was online for about an hour, it was installed on 15 machines. This included a case where a security company's scanner executed the code, resulting in the theft of credentials.
In another case, after scanning approximately 9,000 targets, Claude successfully infiltrated a company's application by combining publicly available credentials with SQL injection (an attack technique that injects malicious commands into a database).
A Statement of a Corporate Culture of "Blameless Postmortem"
In its blog post, Anthropic stated, "While many factors were involved in this incident, we will adhere to our culture of blameless postmortem and take full responsibility for the corrections." This clearly demonstrates their refusal to shift blame, even though the direct trigger was a configuration error on the part of their evaluation partner.
Anthropic stated that they notified the affected organizations on Monday (July 27th), and that both organizations were unaware of their breaches until Anthropic informed them.
What Engineers Should Consider
The two incidents involving OpenAI and Anthropic highlight the structural difficulties inherent in the security evaluation of AI agents. The evaluation method of "measuring the true capabilities of a model with safety features removed" is necessary to understand the model's limitations, but it also carries the risk of real-world consequences if there are vulnerabilities in the evaluation environment itself.
In this case, the cause was not a novel vulnerability, but rather a long-standing, basic security flaw such as "weak passwords" and "endpoints that do not require authentication." This series of incidents illustrates that when designing systems that give AI agents broad network access, more attention should be paid to the surrounding environment and the verification of access boundaries than to the capabilities of the model itself.