Thursday, September 10, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Morning Edition Free of Charge
TODAY'S LEAD STORY

When you try to get AI to "create an AI agent," only 23.9% succeed—a meta-level evaluation reveals a new barrier in the industry.

This article explains "τ^τ-Bench," which Sierra released as open source on September 8th. It describes a 53-task, 4-domain design for building a customer service agent that operates under the same conditions as actual outsourcing, including scattered business records, conversational clients, production APIs, and existing codebases. It highlights the significant difference between the highest-performing Claude Opus 5 + Claude Code (23.9%) and the human + AI standard (82.2%), the score distribution across 6 configurations including GPT-5.6 and Kimi K3 (14.9-23.9%), common failure patterns with human developers such as shallow queries, low communication, and insufficient design comparison, and organizes the research methodology from contamination-resistant design through atomic fact decomposition.

When you try to get AI to "create an AI agent," only 23.9% succeed—a meta-level evaluation reveals a new barrier in the industry.
(Photo: illustrative)

AI Only Succeeds 23.9% When Asked to "Create an AI Agent"—Meta-Level Evaluation Reveals a New Industry Barrier

On September 8th, Sierra released a benchmark called "τ^τ-Bench (Hyper Tau Bench)" as open source. While traditional benchmarks measured "how well an AI agent itself can perform tasks," this new benchmark measures a meta-level capability: "how accurately an AI agent can build another AI agent from actual business requirements." As a journalist with an AI research background, I want to examine the ingenuity of this evaluation design and the implications of its results.

The Idea of ​​Making "Agent Construction Itself a Task"

The originality of this benchmark lies in its problem setting itself. In recent years, the ability to "use" AI agents for customer service has become commonplace. Sierra's research team points out that the more fundamental question is shifting to "who builds that agent in the first place, and how?"

Furthermore, despite the increasing reliance on AI coding agents for the "creation" process itself, existing benchmarks have largely neglected to evaluate this "construction process itself." τ^τ-Bench is designed to fill this gap.

A Realistic Starting Point: "Scattered, Inconsistent Business Records"

The task design evaluated by this benchmark faithfully replicates the realities of actual consulting work. The development agent is given the same starting point as in actual outsourcing: business records actually held by the company, a client with requirements (interacted with interactively), a production API that operations must pass through, an existing codebase to be inherited, and constraints on serving costs and models.

From there, the agent is required to deliver a fully functional, complete customer service agent. The results are evaluated after deployment to pre-prepared (not used for training) simulation users. It consists of a total of 53 tasks spanning four domains.

A Shocking Difference: 23.9% vs. 82.2%

The results of this benchmark were, frankly, disappointing. Even the best automated configuration—Claude Opus 5 running on Claude Code with maximum inference settings—only passed 23.9% of the evaluation simulation.

In contrast, the expert reference (benchmark) of "a human engineer collaborating with a model of the same level with deep contextual understanding" achieved 82.2% on the same task. In other words, even the best automated agent currently available falls short of one-third of the level achieved by human-AI collaboration.

Practical Reference Information: Rankings of Specific Models

The research team compared multiple automated developer configurations, and the results provide interesting practical reference information. The scores of the six automated configurations ranged from 14.9% to 23.9%. Following the highest-scoring Claude Opus 5 (Claude Code, maximum inference), the Codex running GPT-5.6-Sol with Xhigh inference settings came in at 22.0%, followed by Codex using GPT-5.6-Terra at 18.0%, OpenCode using Kimi K3 at 17.9%, Kimi Code, also using Kimi K3, at 16.1%, and Claude Code using Claude Sonnet 5 at 14.9%.

The time taken to build also varied greatly depending on the configuration. The fastest Codex using GPT-5.6-Terra averaged 30.0 minutes, while the longest-taking OpenCode using Kimi K3 averaged 360.3 minutes (6 hours).

An interesting observation: "The same failure patterns as human developers"

Particularly insightful in this study is the observation that the failure patterns encountered by AI agents are remarkably similar to those actually faced by human agent developers. The model tends to rely on superficial queries (searches and inquiries) instead of deeply understanding business records, engaging in minimal communication with clients, and shipping the initial design without sufficient experimentation with the agent's architecture or serving costs.

This suggests that the limitations of AI coding agents lie not merely in a lack of technical capability, but in a more fundamental aspect of "engineering attitude"—the inability to deeply explore requirements and compare multiple design options.

Methodological Honesty: A Design Robust to Contamination

In this benchmark design, the description of the task being evaluated is broken down into mechanically checkable "atomic facts," and a mechanism is in place to mechanically verify whether the generated deliverables accurately reflect the assigned facts. This design is similar to the concept of "contamination tolerance" adopted by the previously discussed Reconstruction benchmark. The inclusion of a mechanism to rigorously check whether the assigned requirements are actually met, rather than simply generating "plausible" output, enhances the reliability of the evaluation.

What Researchers Should Observe

The lesson that τ^τ-Bench teaches is that there is still a significant gap between "the ability of an AI coding agent to write code" and "the ability of an AI coding agent to understand actual business requirements and deliver a deployable finished product." This relates to the previously discussed "model fatigue," but while the performance competition among individual models intensifies, the limitations of capabilities revealed by this "meta-level evaluation simulating actual outsourcing" provide an important perspective for evaluating the practicality of AI agents that cannot be captured by simple coding benchmarks alone.

Going forward, we will continue to closely monitor how this benchmark score evolves with the generational shift of models, and how close automated agents can get to the 82.2% human collaboration standard.

ベンチマークAIエージェントAI/ML研究方法論コーディングエージェント

The figure of "$5.6 million" was a lie—NSA, CISA, and FBI expose the systematic model extraction of six Chinese companies.

The joint recommendation AA26-251A, issued on September 8th by the NSA, CISA, and FBI, accuses six companies—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—of systematic knowledge distillation of US frontier models. This article explains the criticism that DeepSeek's "$5.6 million" training cost is a misleading figure that does not include the true cost of the distilled data, names 41 models including Claude/GPT/Gemini/Grok and maps them to MITRE ATLAS10 methods, explains Moonshot AI's extraction of Claude Fable, the detection evasion technique called "transfer stations," and recommends countermeasures to quietly weaken models.

The "$5.6 Million" Figure Was a Lie—NSA, CISA, and FBI Expose the Systematic Model Extraction by Six Chinese Companies

On September 8th, the U.S. National Security Agency (NSA), the Cybersecurity and Infrastructure Security Agency (CISA), and the FBI jointly issued cybersecurity advisory "AA26-251A." It specifically accuses six Chinese companies—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—of conducting a systematic, industry-scale "knowledge distillation" campaign against U.S. frontier AI models since the end of 2024. As an engineer, I want to examine the specific technical methods outlined in this advisory and their implications.

An Unusually Detailed Document with "41 Models" and "10 Methods"

First, it's important to understand the nature of this advisory. It's unusually detailed for this type of advisory issued by government agencies. The recommendation names 41 versions of US-made AI models and maps their behavior to 10 MITRE ATLAS techniques (an industry-standard framework that systematizes attack methods against AI systems). Furthermore, it describes four new tactics not yet covered by the existing MITRE ATLAS framework.

The technique of distillation itself is not inherently illegal. The method of training a smaller, newer model using the output of a larger, high-performance model is a legitimate research technique commonly used in all major AI labs (including US companies), provided it is authorized. The recommendation is not concerned with the technique itself, but with the unauthorized and large-scale "extraction" conducted in violation of each company's terms of service.

The Misleading Figure of "$5.6 Million"

The most specific and shocking point in this recommendation concerns DeepSeek's training cost. The recommendation points out that DeepSeek's publicly advertised low model training cost of "$5.6 million" is actually misleading. These figures do not include the true cost of data obtained through malicious distillation.

According to the recommendations, DeepSeek allegedly extracted specialized training data and capabilities from a wide range of US-made models—including Claude 3.7, Claude Sonnet 4, Claude Sonnet 4.5, Claude Opus 4.1, Gemini 2.5 Pro Preview, Gemini 2.5 Flash Preview, GPT-4, GPT-4o, GPT-4 Mini, GPT-4 Nano, GPT-5, and Grok 4—between the end of 2024 and mid-2025, and used this data to train its own R1 and V3 models.

Extraction from Claude Fable by Moonshot AI

Specific correspondences are also shown for other companies. Moonshot AI is said to have conducted a widespread distillation campaign since mid-2025, particularly by extracting large amounts of data from Anthropic's Claude Fable 5 and using it to train its "Kimi-K3" model. It is also alleged that they developed "Kimi K2" using GPT-4o output.

Alibaba is said to have leveraged industrial-scale distillation to improve its "Qwen" family of models. The recommendation assesses that all six companies were likely conducting these activities "with the knowledge of the Chinese government."

The "Transfer Station" Method of Evading Detection

A technically interesting method revealed in this recommendation is the use of a "transfer station," a bypass route involving multiple API proxies. China-based AI companies are alleged to have used this type of route to circumvent regional restrictions and gain unauthorized access.

According to International Business Times, this method included the use of fraudulent accounts, the acquisition of large numbers of premium subscriptions, and the use of proxy routing services. This suggests that the attacks were not isolated incidents, but rather part of a systematic and continuous operational structure.

An Unusual Countermeasure: "Silently Weakening the Model"

Another noteworthy aspect of this recommendation is the content of its "countermeasures." CISA recommends that U.S. AI providers, instead of simply banning access to suspicious accounts, adopt a silent quality degradation approach: "providing a weaker model without notifying the account of the issue."

The indicators listed as signs to detect include behavioral metrics such as "continuous 24/7 usage without human-like fluctuations or downtime," "new contracts reaching maximum usage immediately after signing, rather than phased deployment," "shared accounts spanning multiple IPs and user agents," and "unusual ratios of subscription volume to API usage." A distinctive feature is the emphasis on analyzing usage patterns rather than traditional forensic evidence such as hash values ​​or IP blocklists.

A Reservation Regarding "Attribution Claims"

However, one important reservation is necessary when evaluating this recommendation. The contents of this recommendation are merely "claims" by a U.S. government agency, and no independent third-party verification to support them has been conducted at this time. As repeatedly pointed out in our previous analysis of cyberattacks, this type of attribution issue involves extremely difficult technical judgments.

Treasury Secretary Scott Bescent has reportedly threatened sanctions and designation on the Entity List (a list of companies subject to export restrictions) against companies distilling U.S. models. It will be important to closely monitor how the counterarguments from the named companies and independent verifications regarding the contents of this recommendation develop in the future.

What Engineers Should Consider

This recommendation demonstrates the reality that the contractual constraints of AI model terms of use, which have often been relatively overlooked, are now being elevated to a national security concern. For businesses providing their AI APIs, this type of "behavioral-based anomaly detection" method could serve as a practical guideline for future misuse prevention measures.

On the other hand, for companies that utilize AI models, the need to verify the origin of the models they use and whether there are any such suspicions regarding the model's learning process will likely increase as part of the procurement process.

サイバーセキュリティAI安全性DeepSeek国家安全保障AI/ML

A factory where "robots make robots" has begun operations—XPeng's IRON mass production line widens the gap with the Tesla Optimus.

On September 8th, XPeng announced the launch of a mass production line for its IRON humanoid robot in Guangzhou, featuring over 80% automation of core processes and a design built from scratch. This article examines the difference between "launch" and "start of mass production" (full-scale mass production is expected by the end of 2026, with overseas expansion planned for 2027), the contrasting approach with Tesla Optimus, which is currently modifying its existing production line, the reconsideration of Musk's January statement (that no useful work was yet done), the reaffirmation of the previously announced specifications of 76 degrees of freedom and 2,250 TOPS, and the accounting status of the project, which will remain consolidated even after raising $900 million.

A Factory Where "Robots Build Robots" Begins Operation – XPeng's IRON Mass Production Line Widens the Gap with Tesla Optimus

On September 8th, XPeng announced the official launch of its mass production line for the humanoid robot "IRON" at its Guangzhou factory. This announcement, accompanied by footage of a fully assembled IRON robot walking off the production line without human assistance, is a concrete progress report following the previously reported $900 million funding round. As a software engineer, I want to examine the details and limitations of this "mass production line launch."

Specific Figures: "80% of Core Processes Automated"

First, let's look at the technical details. XPeng claims that this production line achieves over 80% automation of core processes. The company positions this as "the world's first automated production line for advanced humanoid robots," explaining that it applies the automotive-grade quality control systems cultivated in its automotive business to the new field of robot manufacturing.

CEO He Xiaopeng stated, "This robot production line is unprecedented and built from scratch." The fact that it was newly designed to suit the unique assembly processes of humanoid robots, rather than repurposing an existing automobile production line, is in stark contrast to Tesla's approach of modifying existing Model S/X production lines for Optimus production, as previously discussed.

A Precise Understanding: "Mass Production Line Operation" Not "Mass Production Begins"

It's crucial to understand the context of this announcement precisely. As XPeng itself has stated, this "line operation" does not yet mean the "start of mass production." Full-scale mass production is planned to begin by the end of 2026, with initial deployment limited to its own stores and campuses. Formal launch and delivery in China and overseas are expected in 2027.

In short, this milestone represents an important transition from "prototypes in the R&D stage to actual production line manufacturing," but it does not yet represent "mass commercial shipments." This distinction serves as a reminder of the lesson we previously discussed: the gap between announcement and verified performance.

A Striking Contrast with Tesla Optimus

One reason this announcement is attracting even more attention within the industry is the contrast with Tesla Optimus. Musk had previously stated that he expected to produce 10,000 Optimus vehicles by 2026, but in January, he admitted that none of the robots were actually performing any useful tasks yet. As previously discussed, the conversion of the Fremont plant to the Optimus production line has been delayed from its original schedule, and the pace of production is expected to be "very slow."

As Electrek's analysis aptly points out, the fact that "Tesla is still in the process of converting its car production lines, while XPeng has already got its lines up and running" symbolically illustrates the difference in progress between the two companies. However, it should also be emphasized that "one robot walking off the line" does not yet mean "mass production" itself.

Reconfirming Already Seen Specifications: "76 Degrees of Freedom" and "2,250 TOPS"

The basic specifications of IRON remain unchanged from our previous article on fundraising. Its features—76 degrees of freedom throughout the body, 21 degrees of freedom in each arm, a maximum computing performance of 2,250 TOPS thanks to three Turing AI chips, and a flexible grid-like covering structure—remain the same. This configuration allows the Physical AI platform model to operate directly on the robot body, enabling it to autonomously perform complex tasks without relying on remote control.

Accounting Positioning: "Consolidation" in Financial Statements

We would also like to reconfirm another important point that was revealed in the August fundraising reports. Even after raising over $900 million in external funding, XPeng maintains a controlling stake in its robotics business, and this business continues to be included in the group's consolidated financial statements. This means that the profit and loss of the robotics business will continue to have a direct impact on XPeng's overall performance.

What Software Engineers Should Observe

The launch of the IRON mass production line is a significant step forward in the humanoid robot industry, demonstrating a steady shift from "research and development" to "the construction of actual manufacturing infrastructure." In particular, the approach of directly applying the quality control system cultivated in existing automobile manufacturing to robot production is a sound strategy from the perspective of mass production reproducibility.

However, there is still a significant gap between "the line is operational" and "actually being able to mass-produce thousands or tens of thousands of units with high reliability." Considering He Xiaopeng's own statement, previously mentioned, that "developing humanoid robots is 20 times more difficult than developing automobiles," it should be assumed that many technical challenges await beyond this milestone. We will be closely watching the actual mass production scheduled for the end of 2026 and its scale.

XPengヒューマノイド製造業TeslaフィジカルAI
Advertisement300 × 250