Wednesday, September 9, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Morning Edition Free of Charge
TODAY'S LEAD STORY

The price of "14 times fewer tokens"—concerns that OpenAI's new model makes it impossible to read thoughts.

This article explains the security controversy surrounding OpenAI Astra's adoption of "recursive depth (loop-type transformer)," as reported by The Information on September 2nd. It covers the mechanism that achieves high performance with up to 1/14th the number of tokens compared to conventional methods by repeatedly passing tokens through the same block, concerns about weakened thought chain monitoring, Chief Scientist Jakub Pachocki's rebuttal four hours later and his explanation of the computation graph depth being within twice that of GPT-4, the ambiguous conclusion that "new rallies are not used, but monitoring capabilities are reduced," and the need for regulatory and auditing systems pointed out by Dean Ball, all from a technical standpoint.

The price of "14 times fewer tokens"—concerns that OpenAI's new model makes it impossible to read thoughts.
(Photo: illustrative)

The Price of "14 Times Fewer Tokens"—Concerns About Unreadable Thoughts Caused by OpenAI's New Model

On September 2nd, a report by The Information sparked a heated debate within the AI ​​security community. The report concerned OpenAI's next-generation frontier model, "Astra," which employs a new architecture called "recurrent depth," or "loop transformer," potentially making the model's "chain of thought" unreadable to humans. As an engineer, I want to examine this technical mechanism and the resolution of the surrounding debate.

A New Design: "Passing Through the Same Block Multiple Times"

First, let's understand the mechanism of this "loop transformer" technology. A typical transformer model passes input tokens (fragments of words) through each layer of the network only once in sequence. However, in a loop design, tokens are repeatedly passed through a single block (a part of the network that can consist of multiple layers) multiple times. Furthermore, the output of that block is not recorded in a "scratchpad" (an area where human-readable text is written) each time, but is instead sent back to the same block.

The advantage of this mechanism lies in its ability to extract higher performance relative to the model size, because it repeatedly uses the same mathematical processing and does not require passing all tokens through all layers of the network. According to reports, Astra achieves superior results with up to 1/14th the number of output tokens compared to conventional models using this method.

Impact on the "Chain of Thought" Monitoring Method

The reason this technology is causing concern among security researchers is that it could weaken one of the important methods for ensuring the security of AI. "Chain of Thought monitoring" is a method where, before an AI model reaches a final answer, its reasoning process is written out as natural language text, allowing humans to read it and externally verify what the model thought and why it reached that conclusion.

In loop-type transformers, part of this reasoning proceeds not as natural language text, but as a numerical representation (sometimes called "neuralalis") processed within blocks. This numerical representation is far more information-dense than human language, yet completely incomprehensible to humans. AI safety researcher Ryan Greenblatt expressed strong concern, stating that if the report were true, it "could be one of the worst developments in AI safety to date."

OpenAI Chief Scientist's Quick Damage Control

Just four hours after this report, OpenAI's Chief Scientist, Jakub Pachodski, posted a rebuttal on X (formerly Twitter). He explicitly stated that "Astra does not use neuroalalis," and provided technical specifics, explaining that "the computational graph depth of current frontier models (including Astra) is within twice that of GPT-4."

Pachotsky further stated, "OpenAI has been working to maintain and utilize thought chain monitoring since its first inference model. This monitoring is vulnerable and, unfortunately, is tending to worsen, but this is not due to changes in architecture." He added that efforts to strengthen this technology are a core goal of the current research program.

A Subtle Conclusion: Not a Complete Denial

Multiple analyses that carefully followed this exchange point out that OpenAI's response was not a "complete denial." While Pachotsky stated that Astra "does not use" Neuralis, he did acknowledge the fact that "monitorability has decreased."

In other words, while the reported, more extreme characterization of "Neuralis" is inaccurate, OpenAI itself did not deny the more moderate fact that Astra's new architecture somehow weakens the effectiveness of the thought chain monitoring safety tool.

Questions about Governance in a "Twitter Settlement"

AI policy commentator Dean Ball, who observed this series of events, makes an interesting point. He argues that the panic surrounding the false claims about "New Rallies" paradoxically highlights the need for regulation, particularly the rapid institutionalization of audits and technical assessments of frontier AI labs.

Ball's observation that "we are now judging technically complex and nuanced claims at the speed of a Twitter timeline, with almost no grounded information on what is actually happening" points to a deeper issue regarding transparency in this field.

What Engineers Should Consider

This incident demonstrates the tension between two goals: improving the efficiency of AI models and maintaining their transparency and monitorability. The dramatic 14x token reduction is a highly attractive improvement in terms of reducing computational costs. However, if that efficiency is achieved at the expense of ease of human monitoring, then it is necessary to fully understand the cost before deciding whether or not to adopt it.

How will the technical measures to prevent the internal workings of AI models from becoming a complete black box—what Pahotsky calls "enhancement efforts"—be concretized in the future? And will this kind of technical debate be resolved through a more systematic verification process rather than through exchanges on Twitter? We will be watching future developments closely.

OpenAIAI安全性TransformerAI/MLモデルアーキテクチャ

What the number "3.1 agent working days" means—Internal data for self-improvement released by OpenAI

This article explains "Research acceleration: The view inside OpenAI," released by OpenAI on September 6th. It reports on the achievement of the goal of "automated research interns by September 2026," which was set in the fall of 2025; the ratio of 3.1 agent working days per human working day as of mid-August; the median inference cost of $600 per researcher per day, with the top 10% exceeding $7,000; the measurement method using a six-stage classification system for Epoch AI (decision, design, build, execute, analyze, and communicate); a self-warning that "increased workload does not equal increased research output"; and the next goal of becoming an automated AI researcher by March 2028, all organized as research methodology.

What the Number "3.1 Agent Working Days" Means—OpenAI's Internal Data on Self-Improvement

On September 6th, OpenAI released a report titled "Research Acceleration: The View Inside OpenAI." This report reveals internal data showing how much their coding agents are accelerating the company's own research activities. As a journalist with a background in AI research, I want to examine the technical details and implications of this data showing "AI accelerating its own research."

The Achieved Goal: "Automated Research Interns"

First, it's important to understand the report's nature as a report on goal achievement. In the fall of 2025, OpenAI had already publicly announced its goal of "achieving automated research interns by September 2026." This report is a progress report on this goal, and according to OpenAI's own measurements, this goal has already been achieved.

The definition of "research intern" here is crucial. OpenAI describes this as "a system capable of performing clearly defined research tasks under human direction, including tasks that would take a skilled researcher several days." This emphasizes the difference in nature: it's not an "autonomous scientist" that autonomously chooses research topics and pursues them independently, but rather a "competent agent" that performs specific tasks under human supervision.

The Precise Meaning of the Number "3.1"

The most attention-grabbing aspect of this report is its unique metric, "agent-workday." As of mid-August, OpenAI's research organization was performing 3.1 agent-workdays' worth of work for every human-workday. This ratio was still lower than human work before June 2026, meaning this reversal occurred relatively recently.

However, caution is needed when interpreting this number. As OpenAI itself clearly warns, this "3.1 times" figure does not simply mean that "researchers have become 3.1 times more productive." The agent's operating time can be parallel, redundant, unsuccessful, or require strong human guidance, meaning research progress doesn't necessarily correlate directly with this raw activity level.

The Reality of Computational Costs: From $600 to $7,000 per Day

Another interesting aspect of this report is the disclosure of specific computational costs. As of mid-August, the median researcher was spending over $600 per day on inference costs for internal agents (in API price equivalent). Furthermore, the top 10% of users were reportedly consuming tokens exceeding $7,000 per day.

These figures demonstrate that using AI agents in research is no longer a negligible cost, but is becoming a substantial expenditure item that should be clearly considered in the budget allocation of research organizations as a whole.

The Ingenious Measurement Method: "Epoch AI's Six-Stage Classification"

Another noteworthy aspect of this report is the measurement method itself. OpenAI classified the token usage of coding agents using a classification method published by Epoch AI (an independent organization that analyzes the progress of AI research), which divides the AI ​​research and development lifecycle into six stages (decision, design, build, execute, analyze, and communicate).

By using this structured classification method, OpenAI aims to gain a more precise understanding of how much each agent contributes to each stage of the research process. The fact that this analysis goes beyond mere aggregation of "usage" and adheres to the structure of the research process itself demonstrates the academic integrity of this report.

The Broader Context of "Recursive Self-Improvement"

OpenAI positions this data release as part of a broader context—"recursive self-improvement," the phenomenon where AI accelerates its own development. The company states that this will be one of the most important AI developments in the coming years and emphasizes the importance of building a "shared public understanding" of how this progress is occurring.

What's interesting is that this report was published in the same week as the "Astra Neuralis controversy" that we previously discussed. The simultaneous pursuit of dramatic efficiency improvements and efforts to increase transparency in AI-driven research acceleration seem to symbolize the ongoing tug-of-war across the entire industry between two goals: "speed" and "explainability."

The Next Goal: March 2028

In this report, OpenAI reiterates its next goal: to realize more advanced "automated AI researchers" by March 2028. This suggests a gradual transition from the current "interns performing tasks under human supervision" to a more autonomous system capable of participating in research topic selection and progress.

What Researchers Should Consider

The biggest lesson this report offers is that the progress of AI research is no longer solely the domain of human researchers; its very speed is changing due to collaboration with AI agents. However, as OpenAI itself repeatedly emphasizes, an increase in agent operating time does not automatically equate to an increase in meaningful research results.

How can we bridge this gap between quantity and quality? And to what extent can the acceleration of AI-driven research be verified externally? We will continue to closely monitor the progress of transparency in this field, including whether other frontier research institutes will release similar internal data in the future.

OpenAIAI研究AIエージェント自己改善AI/ML

1.8 trillion yen without owning a single chip—FluidStack demonstrates a new way to make money with AI infrastructure.

On September 5th, FluidStack reportedly reached a valuation of $18 billion with a $1.5 billion funding round led by Jane Street. This article will analyze, from an accounting perspective, the rapid growth from $1.8 million in 2022 to a projected $660 million in 2026, its "chip-independent" neo-cloud business model (without owning its own GPUs), the contrast with CoreWeave's $28.5 billion hardware funding round, the relocation of its headquarters from the UK to the US and withdrawal from France due to its $50 billion contract with Anthropic, the correlation with Jane Street's own quarterly $16.1 billion in AI transaction revenue, and the risks of dependence on a single giant customer.

1.8 Trillion Yen Without Owning a Single Chip—FluidStack's New Way to Profit from AI Infrastructure

FluidStack, a startup specializing in AI data centers, completed a $1.5 billion funding round led by Jane Street, reaching a valuation of $18 billion, as reported on September 5th. What's remarkable about the company is that it has built such a massive business scale, commensurate with this valuation, without owning a single GPU (high-performance chip used for AI processing). As an accountant, I want to examine the structure of this business model.

The Steep Growth Curve: From $1.8 Million to $660 Million

First, let's look at the company's growth pace in numbers. FluidStack's revenue was reportedly only $1.8 million in 2022. This grew to $66.2 million in 2024, and is projected to reach approximately $660 million by the full year 2026. In just four years, their revenue has ballooned to more than 360 times its original size.

This rapid growth is supported by the $50 billion contract with Anthropic, which we previously discussed. Under this contract, FluidStack is building custom-designed data centers in Texas and New York to meet Anthropic's needs. While the company was previously known for providing infrastructure to Mistral, this massive contract propelled them to become a major player in the AI ​​infrastructure industry.

A Unique Position in the Industry: "No Chip Ownership"

The most interesting aspect of this business model is the fact that FluidStack itself does not own GPUs. The company positions itself as a "chip-agnostic" AI neo-cloud (an AI-focused cloud infrastructure provider). In other words, it specializes in the upstream expertise of data center design, construction, and operation, without bearing the risks of hardware procurement and ownership.

This positioning is a stark contrast to competitors like CoreWeave, which we previously discussed, that own a large number of GPUs in-house. While CoreWeave raised $28.5 billion in equity and debt over the past year and plans $30 billion to $35 billion in capital expenditures for 2026, FluidStack is growing its business without including such massive hardware procurement costs on its balance sheet.

Shifting Headquarters from "UK-based" to "US-based"

Another interesting transformation of FluidStack is its relocation. Originally a spin-off from Oxford University, the company attracted attention as a rising star in the European AI scene. However, due to the importance of its large contract with Anthropic, the company moved its headquarters from the UK to the US.

Furthermore, it has been reported that in March of this year, the company completely withdrew from its French projects and shifted its focus to the US AI infrastructure market. This series of developments is a symbolic example of how strongly the focus of AI infrastructure investment is being drawn to the US market.

The "Investor Profile" Shows the High Valuation of This Deal

Jane Street, which led this fundraising round, is a company known for quantitative trading (financial trading using mathematical methods), and is said to have generated enormous profits from AI-powered trading strategies, reaching $16.1 billion in the first quarter of 2026 alone. The fact that a player with such a profound understanding of computing power in the financial world is leading investments in an AI infrastructure company is an intriguing coincidence.

It was also reported that Google was considering investing in FluidStack during previous fundraising negotiations. This is likely related to Google's strategy of more broadly deploying its TPUs (computer-aided processing units, proprietary chips for AI) through third-party data center operators.

Positioning in the Growing "$350,000 Neo-Cloud Market"

The "neo-cloud" market segment to which FluidStack belongs is projected by Synergy Research to reach approximately $400 billion by 2031, representing an extremely rapid growth rate of 58% per year. Within this market, FluidStack is establishing itself as a leading independent player, second only to the publicly traded company CoreWeave.

Points to Note from an Accountant's Perspective

FluidStack's business model demonstrates an interesting precedent: amidst the AI ​​infrastructure investment boom, it's possible to build enormous corporate value solely through expertise in construction and operation, without taking on the risk of owning hardware. While GPUs themselves are susceptible to rapid obsolescence due to technological advancements, expertise in data center design and construction is a more robust asset against such technological generational shifts.

FluidStack's business model demonstrates that, amidst the AI ​​infrastructure investment boom, it's possible to build enormous corporate value solely through expertise in construction and operation, without taking on the risk of owning hardware. On the other hand, the company's valuation is heavily dependent on its contract with a single, massive customer, Anthropic. The financial situation of Anthropic itself, as discussed previously, and its moves toward an IPO could directly impact FluidStack's corporate value. This is a crucial risk factor to consider when evaluating the company. Going forward, how many other major customers FluidStack can secure will likely be a key indicator of the sustainability of its valuation.

FluidStackAIインフラデータセンター資金調達ファイナンス
Advertisement300 × 250