Monday, September 21, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Morning Edition Free of Charge
TODAY'S LEAD STORY

The day "folders" became "coordinators"—Understanding the redesign of Claude Code Projects

On September 17th, Anthropic completely revamped the Projects feature of Claude Code, releasing a beta version of a system where a single conversation acts as a coordinator, managing parallel cloud threads. This article explores the new design philosophy of AI coding tools, including the two-tiered structure of coordinators and workers, the differentiated use of models and effort for each role, and the usage limit of 200 threads per day.

The day "folders" became "coordinators"—Understanding the redesign of Claude Code Projects
(Photo: illustrative)

The Day "Folders" Became "Coordinators"

On September 17th, Anthropic completely redesigned and released a beta version of Claude Code's "Projects" feature. Previously, Projects were like "folders," bundling several files and a single chat. The new Projects system is reborn, with a single ongoing conversation acting as a coordinator, breaking down larger goals into multiple parallel threads for execution. The title of the announcement, "From Folders to Conversations," succinctly expresses the essence of this change.

A Two-Tier Structure: "Coordinator" and "Workers"

From a technical design perspective, the two-tier architecture of project conversations and threads is interesting. The project conversation acts as the command center, interpreting user instructions, answering simple questions on the spot, and creating threads for the parts requiring actual work. Each thread is an independent Claude Code cloud session, each with its own dedicated branch and repository copy. The thread opens a pull request, runs tests, reads the documentation, and reports a summary of the progress to the coordinator. The coordinator doesn't see every step of the thread; they only receive the reported summary. This design makes sense in preventing context bloat.

Practical Examples

The examples accompanying the presentation are easy to understand. If you want to reduce the p75 latency (the time it takes for 75% of requests to receive a response) of the checkout process, Claude profiles each endpoint and opens a pull request while trying multiple optimizations in parallel threads. Alternatively, if you connect multiple API, Web, and Mobile repositories and give the goal of "retiring deprecated v1 endpoints," Claude creates a thread for each repository, handles caller migration, test execution, pull request creation, and even reports which pull requests should be merged first. The ability to parallelize the tedious but time-consuming migration work across multiple repositories has a significant practical impact.

Differentiating Models and Effort Levels by Role

Another practical design consideration is the ability to individually set the model and effort level (thought level) used for both coordinators and worker threads. This allows for flexible resource allocation; for example, assigning the highest-level model and high effort to coordinators handling complex tasks like situation assessment and task decomposition, and assigning the standard model and low effort to workers handling routine test execution and single-function refactoring. There are reports that by default, each thread is set to run Opus at high effort, which can be seen as an engineering technique to reduce token consumption while maintaining quality.

A Practical Caution: Faster Usage Consumption

A practical caution that cannot be overlooked is that parallel threads consume the plan's usage limit faster than usual. There is a limit of 200 new threads per day, and without paid additional credits, threads that reach the plan limit will wait until it is reset. Threads started by routines are an exception; if they reach the limit, they will terminate with an error in that turn, and the user will need to send a message again after the reset. The parallelization-driven productivity improvements come at the cost of increased operational challenges, particularly in managing usage limits.

Phased Rollout Plan

Currently, the beta is limited to a select group of Pro and Max subscribers using cloud sessions, targeting users without existing projects on chat or Cowork. The plan is to expand to Pro and Max users over the next few weeks, followed by Team and Enterprise Plan users. Existing projects will continue to function as usual until the rollout is complete.

What Engineers Should Note

From single-prompt responses to collaborative work involving multiple agents managed by a coordinator—this shift indicates that AI coding tools are moving from "question and answer" to "autonomous project-based progress management." This makes it an attractive option for tasks where parallel exploration is beneficial, such as migration across multiple repositories or performance tuning. On the other hand, the speed at which usage limits are consumed, and the risk of oversights due to coordinators not having a detailed understanding of the thread's progress, will become clearer as the beta expands, revealing how much it can be relied upon in actual operation.

AnthropicClaude CodeAIエージェント開発者ツールクラウドインフラ

From "thought experiments" to "internal dashboards"—Anthropic's R&D automation index measures the speed of AI-driven AI development.

From September 17th to 18th, Anthropic published its "R&D Automation Index," based on Epoch AI metrics, revealing that Claude has reached a level where he "leads" 26% of the company's AI research and development work. This article analyzes the measurement method, which divides approximately 15,000 tasks into 378 categories, the monitoring of 30,000 internal agents, the recursive structure in which Claude himself handles the measurement, and its limitations.

From "Thought Experiment" to "Internal Dashboard"

Recursive self-improvement—the phenomenon where AI accelerates the development of its successor models, and the improvement cycle becomes even faster—has long been a thought experiment debated among AI safety researchers. Between September 17th and 18th, Anthropic provided a concrete measurement for this debate. They announced that the percentage of their R&D work "led" by Claude had risen from less than 1% in February to 26% in August. This announcement came just days after CEO Dario Amodi called on the industry to slow down the pace of development.

The "R&D Automation Index" as a Measurement Method

At the core of the framework announced is the "R&D Automation Index," based on a metric developed by Epoch AI. Anthropic analyzed its internal AI research and development activities, classifying approximately 15,000 tasks into 378 sub-categorization trees. They then evaluated the automation level of each task on a six-point scale, from AL0 (no AI involvement) to AL5 (completed solely by AI without human intervention). As of August, 26% of tasks were judged to be at a level where Claude was "leading"—capable of completing most tasks from high-level prompts, with humans acting as supervisors (AL4). Over 90% of tasks involved Claude playing a role beyond that of a collaborator (AL3 or higher). However, Anthropic explicitly states that complete autonomy (AL5) has not yet been achieved in any of the measured R&D areas.

The Scale of 30,000 Agents

Another indicator is the monitoring system for AI agents operating within the company. Anthropic states that approximately 30,000 agents are constantly running research and engineering tasks on its most widely used agent platform. All agent actions are monitored both in real time and retrospectively, and the extent to which this monitoring network functions is one of the elements that this framework aims to measure. The third metric is the allocation of computing resources. In a sample from July 13th to 20th, approximately 6% of computing resources for AI research and development were allocated to safety research, and this figure rose to approximately 12% when limited to R&D tasks performed by the AI ​​itself. Anthropic also explains that these figures are "conservative," stating that tasks contributing to both safety and capability improvement are not included in the calculations.

The Recursive Structure of Who Did the Measurement

What is interesting from an academic perspective is the recursive structure in which Claude itself is responsible for part of the measurement process. The Claude agent extracts a list of tasks from Slack and internal documents, and another Claude judgment model then assigns an automation level to each individual task. Anthropic itself acknowledges the limitations of this method, pointing out the risk that the judgment model may repeat the same types of errors as the model being evaluated, and emphasizing the need for external verification. Given the self-reported and self-assessed nature of this method, it's important to note that these figures are merely "Anthropic's own measurement of itself" and have not undergone an independent third-party audit.

Proposal to Open Up to Other Companies

Anthropic positions these three metrics as a "reproducible framework that other frontier development companies can adopt." The aim is to advance governance discussions regarding the speed of AI development by accumulating comparable data across industries, rather than having only one company claim its own progress. The company has also revealed plans to actually invite external evaluators into the company and grant them access rights equivalent to those of the risk management team. Acknowledging the limitations of self-reporting and proposing mechanisms to compensate for them deserves some credit as a commitment to transparency.

Anthropic positions these three metrics as a "reproducible framework that can be adopted by other frontier development companies." ## What Researchers Should Note

More important than the 26% figure itself is the fact that a quantitative measurement framework has been provided for the first time to the phenomenon of "AI accelerating AI development," which had previously only been discussed qualitatively. It remains to be seen to what extent this framework will be adopted by other frontier labs, and whether it will eventually be used as the basis for government regulation. However, Anthropic's suggestion that "how this definition is handled may be more important than the number itself" is spot-on. We should keep a close eye on how far independent evaluators will verify the data and whether other companies will publish similar metrics.

AnthropicR&D自動化指数再帰的自己改善AI安全性企業公式発表

What "417 to 3" Didn't Mean: An Accountant's Account of the AI ​​Data Center Electricity Bill's Passage in the House and Rejection in the Senate

On September 16, the U.S. House of Representatives passed the Ratepayer Protection Act by a vote of 417 to 3, requiring AI data centers to bear the full cost of upgrading the power grid. This article examines the performance figures of Oregon's POWER Act, the differences between this act and the White House's voluntary commitment, the structural limitations of the PURPA amendment (which only serves as an "obligation to consider"), and the subsequent rejection of the bill in the Senate.

An Unprecedented 417-to-3 Majority Vote

On September 16th, the U.S. House of Representatives passed the Ratepayer Protection Act, a bill concerning the burden of electricity costs for AI data centers, by an unprecedented 417-to-3 majority. Only three progressive Democrats opposed it, while all Republicans present voted in favor. As an accountant, I was curious about two things: what this bill actually mandates, and why such a bipartisan agreement was reached.

The Contents of the Bill – Not a "Mandatory" Act, but a "Mandatory Consideration" Act

The bill's official name is H.R. 9340, and it amends Section 111(d) of the Public Utilities Regulation and Policy Act (PURPA), enacted in 1978. The bill introduces a "large load standard" for data centers with peak power demand exceeding 100 megawatts (equivalent to the electricity consumption of approximately 80,000 average households). This standard requires data centers to bear the full cost of any additional infrastructure, such as transmission lines, substations, and power generation equipment, built to supply their electricity needs. Even if a contract is terminated, the obligation to bear the cost of already constructed infrastructure remains.

However, as an accountant, I must point out an important caveat. This bill only obligates state utility commissions to "consider" adopting this standard; the decision to actually implement it rests with each state. Ari Pescoe, a power law expert at Harvard Law School, points out that this voluntary promise of price protection does not translate into consumer protection, as state utility commissions still control how costs are actually passed on to household electricity bills.

Figures from Leading States

To predict the effectiveness of this bill, it's helpful to look at the performance of states that have already implemented similar regulations. Oregon's "POWER Act," enacted in 2025, imposes cost burdens on projects exceeding 20 megawatts, resulting in a 30% increase in electricity rates for data centers while residential rates decreased by 1.3%. Virginia also introduced regulations requiring data centers to bear the full cost of power transmission infrastructure starting in July 2026. The fact that electricity rates nationwide have risen by approximately 27% since 2019, outpacing the general consumer price index, explains the widespread bipartisan support for this bill.

Differences from the White House's "Voluntary Pledge"

Interestingly, separate from this bill, major players such as Amazon, Google, Meta, Microsoft, OpenAI, Oracle, and xAI had already signed the White House-led "Ratepayer Protection Pledge" in March of this year. This pledge stipulates that companies must secure their own electricity demand through either "generating, bringing in, or purchasing," bear the costs of the associated transmission infrastructure, and pay fees regardless of whether they actually use the electricity. However, it remains merely a voluntary agreement by companies. The crucial difference between this bill and the current bill is that it attempts to create a legally binding framework. However, as mentioned earlier, even this binding force is only a buffer, requiring "state consideration."

The Senate Outcome—What 417-3 Meaning Didn't

As an accountant, the most interesting aspect is the final fate of this bill. After passing the House of Representatives, the bill was debated in the Senate, but according to reports, it failed to pass due to a block by Democratic senators and was ultimately rejected. Furthermore, the alternative bill, based on a 150-megawatt standard, was similarly blocked. The overwhelming 417-3 vote in the House proved to be no predictor in the face of the different hurdle of unanimous agreement in the Senate. As a result, the authority to allocate costs remains with the FERC (Federal Energy Regulatory Commission) review process and the individual pricing structures for large loads developed by each state's public utility commission.

What Accountants Should Consider

The "417 to 3" figure was a strong signal that the political crisis surrounding the externalities of electricity costs in the AI ​​industry was shared across party lines. However, the effectiveness of the legislation was limited by the structural constraints of the existing federal framework, PURPA, which only mandates consideration, and its failure to pass in the Senate, resulting in less institutionalization than initially anticipated. For data center operators, the practical implications are that, for the time being, the outcome of the FERC's show cause docket and the individually negotiated connection agreements for large loads in each state will continue to be the de facto determinants of cost burden. Rather than waiting for a nationwide federal standard, how to respond to state-level pricing structures remains a realistic challenge for operators.

データセンター電力コスト政策AI規制ファイナンス
Advertisement300 × 250