A Quarter of Tasks Are Wasted
AI coding agents are convenient, but their use comes with a significant financial cost. The paper "Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents," published on arXiv on September 28th, is the first study to systematically measure this "waste behind convenience." A research team from Purdue University and other institutions ran two coding agents, Claude Code and Mini-SWE-Agent, on SWE-bench Verified (a benchmark measuring how many actual GitHub issues can be resolved) under four different configurations, analyzing a total of 1200 execution trajectories (the agent's sequence of actions). The results revealed that cost-inefficient behavior accounted for 79-98% of tasks, representing up to 22.75% of the task cost.
Three "Wasteful Habits"
The paper classifies this waste into three patterns. The first is "subsumed retrieval"—a behavior where the agent re-reads code that has already been retrieved and contains essentially the same information. In Claude Code, this is mainly caused by the parent agent re-reading code retrieved by a sub-agent (an auxiliary agent delegated a task by the parent agent). The second is "sim script generation"—a behavior where the agent repeatedly generates new scripts for similar verification and testing purposes, affecting up to 68% of tasks, and occurring 5.91 to 9.98 times more frequently in Mini-SWE-Agent than in Claude Code. The third is "test re-execution"—a behavior where a test that has already been run is repeatedly executed even though the state has not changed.
Design differences change how waste manifests
What is academically interesting is the discovery that even the same "waste" manifests itself very differently depending on the agent's design philosophy. The guidance (pre-built behavioral guidelines) embedded in Claude Code effectively limited the behavior of sim scripts to primarily "disposable scripts," while the Mini-SWE-Agent showed a strong tendency to repeatedly generate similar files for both testing and editing. In other words, this study quantitatively demonstrates that not only the performance of the agent's underlying model itself, but also how that model is controlled—the "harness"—including prompt design and tool usage constraints, directly impacts cost efficiency.
The Django-13158 Case Study
The paper cites the behavior of Claude Code in the Django repository's SWE-bench challenge "Django-13158" as a concrete example. Through this case study, the research team visualizes how cost-inefficient behavior occurs during actual task execution at the behavioral log level. By including such a concrete case study, the paper encourages readers to understand the causal reasons behind the waste, rather than simply presenting statistical trends.
"Efficiency" in Agent Research: A Previously Understood Axis
Much of the previous research on coding agents has focused on the "resolution rate" of tasks as an outcome metric. This paper points out that, alongside the resolution rate, the axis of "how much wasted cost is involved in that resolution" has not been sufficiently examined. Related prior research includes studies attempting cost reduction through context compression techniques like SWE-Pruner and trajectory shortening, but this paper differs in its approach by focusing on the classification of behavioral patterns themselves—"why this waste occurs."
What Researchers Should Consider
As the practical operational costs of AI coding agents become a realistic consideration among companies, the figure presented in this paper—that "up to 22.75% of task costs are wasted"—has practical implications that go beyond mere academic interest. This research specifically demonstrates that the same resolution rate can be achieved at a lower cost not only by switching to a higher-performance agent base model, but also by reviewing the harness design—mechanisms to prevent duplicate information acquisition, script generation reuse, and test execution state management. Going forward, we will be watching closely to see to what extent this kind of "detection and mitigation of cost-inefficiency patterns" will be incorporated into the evaluation criteria for coding agents themselves.