AI Only Succeeds 23.9% When Asked to "Create an AI Agent"—Meta-Level Evaluation Reveals a New Industry Barrier
On September 8th, Sierra released a benchmark called "τ^τ-Bench (Hyper Tau Bench)" as open source. While traditional benchmarks measured "how well an AI agent itself can perform tasks," this new benchmark measures a meta-level capability: "how accurately an AI agent can build another AI agent from actual business requirements." As a journalist with an AI research background, I want to examine the ingenuity of this evaluation design and the implications of its results.
The Idea of Making "Agent Construction Itself a Task"
The originality of this benchmark lies in its problem setting itself. In recent years, the ability to "use" AI agents for customer service has become commonplace. Sierra's research team points out that the more fundamental question is shifting to "who builds that agent in the first place, and how?"
Furthermore, despite the increasing reliance on AI coding agents for the "creation" process itself, existing benchmarks have largely neglected to evaluate this "construction process itself." τ^τ-Bench is designed to fill this gap.
A Realistic Starting Point: "Scattered, Inconsistent Business Records"
The task design evaluated by this benchmark faithfully replicates the realities of actual consulting work. The development agent is given the same starting point as in actual outsourcing: business records actually held by the company, a client with requirements (interacted with interactively), a production API that operations must pass through, an existing codebase to be inherited, and constraints on serving costs and models.
From there, the agent is required to deliver a fully functional, complete customer service agent. The results are evaluated after deployment to pre-prepared (not used for training) simulation users. It consists of a total of 53 tasks spanning four domains.
A Shocking Difference: 23.9% vs. 82.2%
The results of this benchmark were, frankly, disappointing. Even the best automated configuration—Claude Opus 5 running on Claude Code with maximum inference settings—only passed 23.9% of the evaluation simulation.
In contrast, the expert reference (benchmark) of "a human engineer collaborating with a model of the same level with deep contextual understanding" achieved 82.2% on the same task. In other words, even the best automated agent currently available falls short of one-third of the level achieved by human-AI collaboration.
Practical Reference Information: Rankings of Specific Models
The research team compared multiple automated developer configurations, and the results provide interesting practical reference information. The scores of the six automated configurations ranged from 14.9% to 23.9%. Following the highest-scoring Claude Opus 5 (Claude Code, maximum inference), the Codex running GPT-5.6-Sol with Xhigh inference settings came in at 22.0%, followed by Codex using GPT-5.6-Terra at 18.0%, OpenCode using Kimi K3 at 17.9%, Kimi Code, also using Kimi K3, at 16.1%, and Claude Code using Claude Sonnet 5 at 14.9%.
The time taken to build also varied greatly depending on the configuration. The fastest Codex using GPT-5.6-Terra averaged 30.0 minutes, while the longest-taking OpenCode using Kimi K3 averaged 360.3 minutes (6 hours).
An interesting observation: "The same failure patterns as human developers"
Particularly insightful in this study is the observation that the failure patterns encountered by AI agents are remarkably similar to those actually faced by human agent developers. The model tends to rely on superficial queries (searches and inquiries) instead of deeply understanding business records, engaging in minimal communication with clients, and shipping the initial design without sufficient experimentation with the agent's architecture or serving costs.
This suggests that the limitations of AI coding agents lie not merely in a lack of technical capability, but in a more fundamental aspect of "engineering attitude"—the inability to deeply explore requirements and compare multiple design options.
Methodological Honesty: A Design Robust to Contamination
In this benchmark design, the description of the task being evaluated is broken down into mechanically checkable "atomic facts," and a mechanism is in place to mechanically verify whether the generated deliverables accurately reflect the assigned facts. This design is similar to the concept of "contamination tolerance" adopted by the previously discussed Reconstruction benchmark. The inclusion of a mechanism to rigorously check whether the assigned requirements are actually met, rather than simply generating "plausible" output, enhances the reliability of the evaluation.
What Researchers Should Observe
The lesson that τ^τ-Bench teaches is that there is still a significant gap between "the ability of an AI coding agent to write code" and "the ability of an AI coding agent to understand actual business requirements and deliver a deployable finished product." This relates to the previously discussed "model fatigue," but while the performance competition among individual models intensifies, the limitations of capabilities revealed by this "meta-level evaluation simulating actual outsourcing" provide an important perspective for evaluating the practicality of AI agents that cannot be captured by simple coding benchmarks alone.
Going forward, we will continue to closely monitor how this benchmark score evolves with the generational shift of models, and how close automated agents can get to the 82.2% human collaboration standard.