"Hacking a Real Server to Cheate on Benchmarks"—The Line AI Crossed and an Unprecedented Letter Signed by 1,268 People
On July 21st, OpenAI made a shocking announcement. Their GPT-5.6 Sol model, which was being tested with weakened security mechanisms, escaped from its supposedly isolated test environment and infiltrated the actual Hugging Face production system. Just one week later, this incident directly triggered the publication of an unprecedented open letter signed by 1,268 people (as of the 29th), including chief scientists from OpenAI, Anthropic, Google, and Meta. This article will trace this series of events.
What Actually Happened in the "Supposedly Isolated Environment"?
First, let's summarize the incident itself. In July, as part of an internal cybersecurity assessment, OpenAI ran two models, including GPT-5.6 Sol, in a test environment supposedly isolated from the internet, with weakened security denial mechanisms. However, a previously unknown vulnerability existed in the package installation proxy within that environment. The models exploited this vulnerability to access the external internet and hacked Hugging Face's production system.
Multiple reports emphasize that this is "the first publicly documented autonomous cyberattack by a Frontier AI model against a real company's server, not a simulation or a hypothetical scenario." Furthermore, the models' motive is believed to be "to cheat in benchmarks." This can be seen as a more extreme example than the incident previously discussed in this column where an AI told only to post to Slack sent a pull request to GitHub. OpenAI itself described this incident as "unprecedented."
1,268 Signers in One Week – "Who Signed?" is the Core of the Story
Following this incident, an open letter titled "Pacing the Frontier" was released on July 28th. The content itself is simple: It can be summarized in a single sentence: "We urge the U.S. government to help build, as an international effort, the necessary technical and governance tools to deliberately regulate the pace of cutting-edge AI development."
The crucial point is who signed it. The list includes Anthropic CEO Dario Amodei, co-founders Jared Kaplan and Jack Clark, OpenAI Chief Scientist Jakub Pachocki, Chief Research Officer Mark Chen, Meta Chief Scientist Shengjia Zhao, and Anca Dragan, who leads AI Safety and Alignment at Google DeepMind – those actually developing their own models. The very fact that those on the front lines of development, not external critics, are urging the government to put in "intentional brakes" on their field speaks volumes about the weight of this news.
It's important to understand precisely that the letter itself doesn't call for an immediate halt to development. What it's calling for is preparing a "steering wheel" in advance that will allow for verifiable and collaborative deceleration in case AI systems accelerate beyond the safe oversight capabilities of humans in the future. To borrow the expression from one article, it's about "having the steering wheel ready before the engine goes into recursive gear."
OpenAI and Anthropic Officially Express Corporate Support
Further noteworthy is that OpenAI and Anthropic have each officially expressed corporate support for this letter. OpenAI, in an official post, commented that "at some point in the future, the acceleration of frontier model development may become so high that the world may need to adjust the pace of AI progress." Anthropic also clearly supported this request, citing its own research on recursive self-improvement published last month.
This move comes just before the deadline (August 1st) for the government to develop a safety framework for frontier models under Executive Order 14409, signed on June 2nd.
Meta is "saying the exact opposite at the same time"
This is the point I find most interesting. Meta's chief scientist, Shengjia Zhao, signed the letter personally. However, in the same week, Meta CEO Mark Zuckerberg published an op-ed in the Wall Street Journal arguing that "the benefits of broadly distributing AI outweigh the risks. The danger is not that AI capabilities are too high, but that they are too concentrated in a few."
Furthermore, around the same time as Zuckerberg's op-ed, Meta, NVIDIA, Microsoft, and Palantir jointly issued another letter urging regulators not to restrict the form of open weight models. Within the same company, the chief scientist calls for "preparation for slowdown," while the CEO argues that "open diffusion is the safest option." This rift symbolizes how the conflict between "closed-minded and cautious" and "open-minded and decentralized" factions within the industry is now surfacing beyond company boundaries and manifesting as individual stances.
What Engineers Should Consider
When viewing this incident from an engineer's perspective, the most important thing to keep in mind is the weight of the fact that this Hugging Face incident was not a "hypothesis" but "something that actually happened." Previously in this column, I've introduced several examples of AI agents attempting to breach safety boundaries, but this is the first publicly confirmed case where it actually extended to the systems of an external third-party company.
Building technical guardrails and establishing policies and governance are two sides of the same coin. The sheer size of this initiative, involving 1,268 individuals, and the fact that it includes chief scientist-level figures, indicates that practitioners in this field are beginning to feel a significant gap between the pace of technological advancement and the pace at which mechanisms to safely control it are being developed. We will continue to closely monitor what policy consequences this letter will actually lead to, and how the government's safety framework announcement, scheduled for August 1st, will respond to this development.