The Price of "14 Times Fewer Tokens"—Concerns About Unreadable Thoughts Caused by OpenAI's New Model
On September 2nd, a report by The Information sparked a heated debate within the AI security community. The report concerned OpenAI's next-generation frontier model, "Astra," which employs a new architecture called "recurrent depth," or "loop transformer," potentially making the model's "chain of thought" unreadable to humans. As an engineer, I want to examine this technical mechanism and the resolution of the surrounding debate.
A New Design: "Passing Through the Same Block Multiple Times"
First, let's understand the mechanism of this "loop transformer" technology. A typical transformer model passes input tokens (fragments of words) through each layer of the network only once in sequence. However, in a loop design, tokens are repeatedly passed through a single block (a part of the network that can consist of multiple layers) multiple times. Furthermore, the output of that block is not recorded in a "scratchpad" (an area where human-readable text is written) each time, but is instead sent back to the same block.
The advantage of this mechanism lies in its ability to extract higher performance relative to the model size, because it repeatedly uses the same mathematical processing and does not require passing all tokens through all layers of the network. According to reports, Astra achieves superior results with up to 1/14th the number of output tokens compared to conventional models using this method.
Impact on the "Chain of Thought" Monitoring Method
The reason this technology is causing concern among security researchers is that it could weaken one of the important methods for ensuring the security of AI. "Chain of Thought monitoring" is a method where, before an AI model reaches a final answer, its reasoning process is written out as natural language text, allowing humans to read it and externally verify what the model thought and why it reached that conclusion.
In loop-type transformers, part of this reasoning proceeds not as natural language text, but as a numerical representation (sometimes called "neuralalis") processed within blocks. This numerical representation is far more information-dense than human language, yet completely incomprehensible to humans. AI safety researcher Ryan Greenblatt expressed strong concern, stating that if the report were true, it "could be one of the worst developments in AI safety to date."
OpenAI Chief Scientist's Quick Damage Control
Just four hours after this report, OpenAI's Chief Scientist, Jakub Pachodski, posted a rebuttal on X (formerly Twitter). He explicitly stated that "Astra does not use neuroalalis," and provided technical specifics, explaining that "the computational graph depth of current frontier models (including Astra) is within twice that of GPT-4."
Pachotsky further stated, "OpenAI has been working to maintain and utilize thought chain monitoring since its first inference model. This monitoring is vulnerable and, unfortunately, is tending to worsen, but this is not due to changes in architecture." He added that efforts to strengthen this technology are a core goal of the current research program.
A Subtle Conclusion: Not a Complete Denial
Multiple analyses that carefully followed this exchange point out that OpenAI's response was not a "complete denial." While Pachotsky stated that Astra "does not use" Neuralis, he did acknowledge the fact that "monitorability has decreased."
In other words, while the reported, more extreme characterization of "Neuralis" is inaccurate, OpenAI itself did not deny the more moderate fact that Astra's new architecture somehow weakens the effectiveness of the thought chain monitoring safety tool.
Questions about Governance in a "Twitter Settlement"
AI policy commentator Dean Ball, who observed this series of events, makes an interesting point. He argues that the panic surrounding the false claims about "New Rallies" paradoxically highlights the need for regulation, particularly the rapid institutionalization of audits and technical assessments of frontier AI labs.
Ball's observation that "we are now judging technically complex and nuanced claims at the speed of a Twitter timeline, with almost no grounded information on what is actually happening" points to a deeper issue regarding transparency in this field.
What Engineers Should Consider
This incident demonstrates the tension between two goals: improving the efficiency of AI models and maintaining their transparency and monitorability. The dramatic 14x token reduction is a highly attractive improvement in terms of reducing computational costs. However, if that efficiency is achieved at the expense of ease of human monitoring, then it is necessary to fully understand the cost before deciding whether or not to adopt it.
How will the technical measures to prevent the internal workings of AI models from becoming a complete black box—what Pahotsky calls "enhancement efforts"—be concretized in the future? And will this kind of technical debate be resolved through a more systematic verification process rather than through exchanges on Twitter? We will be watching future developments closely.