Monday, August 24, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Morning Edition Free of Charge
TODAY'S LEAD STORY

Can AI replicate ideas from "unpublished papers"? — The true barrier to scientific reasoning revealed by figures of only 3-15%.

This article explains the new benchmark "Reconstruction," released on arXiv in August. It covers its contamination-tolerant design, which reconstructs the hypotheses of a paper solely from its bibliography; its low scores of 3-15% with single models; the improvement to 42% with a multi-agent approach combining cross-review with multiple models and a Swiss-style tournament, but still not reaching half the target; a survey showing that about half of the 60 benchmarks are saturated with contamination; and the limitations of scientific reasoning in contrast to successful mathematical proofs by AI.

Can AI replicate ideas from "unpublished papers"? — The true barrier to scientific reasoning revealed by figures of only 3-15%.
(Photo: illustrative)

Can AI Recreate Ideas from "Unpublished Papers"? – The Real Barrier of Scientific Reasoning Revealed by a Figure of Only 3-15%

In August, a new benchmark called "Reconstruction" was published on arXiv. It's a unique evaluation method that shows Frontier's large-scale language model only the bibliography of a research paper and asks it to reconstruct what hypotheses the paper actually presented. The results were surprisingly low, with even the best-performing models scoring only 3-15%. As a journalist with a background in AI research, I want to carefully examine what these results mean.

A Clever Design That Guarantees "Not Remembering the Answers"

The technical strength of this benchmark lies in its direct confrontation with a deep-seated problem in AI evaluation: "benchmark contamination." Benchmark contamination refers to the phenomenon where the questions and answers intended for evaluation become embedded in the model's training data, causing the model to essentially "memorize the answers" before taking the test. A saturation survey conducted in 2026 on 60 LLM benchmarks indicated that nearly half of them were potentially already "saturated" by this type of contamination.

Reconstruction employs a design that structurally avoids this problem. If the evaluated paper was published after the model's training data cutoff date, the possibility of its content being included in the model's training data is, in principle, zero. In other words, the model must not "remember" the content of the paper, but purely "infer" what the paper actually argued based solely on the clue of the bibliography.

The Harsh Reality: "3-15% on its own"

The result measured under this clever design is the 3-15% score mentioned at the beginning. This figure indicates that the frontier model's ability to accurately reconstruct the research idea itself from indirect clues such as bibliographies is currently extremely limited.

The research team argues that this result reflects the true capability ceiling the model faces, rather than being an apparent low due to poor prompt design or some flaw in the evaluation method. The figure is compelling precisely because the measurements were taken in a contamination-free environment.

Multi-agent approach improves to 42%—still not quite half

Interestingly, the research doesn't end there. The research team also evaluated the same task using a "multi-agent approach" combining multiple AI models. Specifically, they used a pipeline combining mutual review by multiple models with a "Swiss tournament" (a highly fair tournament format where the next opponent is determined by wins and losses) to select the top four candidate options. They did not use external web searches, relying solely on the provided reference information for inference.

This collaborative approach with multiple models raised the accuracy to 42%. This is a significant improvement compared to the best score of a single model. However, it still remains below half. This suggests that even combining multiple AIs still doesn't reach the level of "expert scientific intuition." ## The Difficulty of Taking These Results "At Face Value"

On the other hand, caution is needed when interpreting these research results. Whether the 15% figure, which achieved the top score, truly constitutes "inference" in a philosophical sense remains debatable. It's impossible to determine from this benchmark score alone whether it simply outputs statistically "plausible" hypotheses from combinations of references, or whether it involves some meaningful reasoning process.

However, the "practical ceiling" demonstrated in this paper—the fact that current models have clear limitations in their ability to reconstruct truly new research ideas from reference clues—can be empirically supported by the robust design of contamination resistance.

The Irony of Timing: Precisely During a Series of High-Profile Announcements

The timing of this paper's publication is also highly suggestive. In August 2026, several Frontier Institutes announced a series of spectacular scientific achievements, such as AI solving unsolved mathematical problems and discovering vulnerabilities in cryptographic algorithms. The mathematical proof by OpenAI's Astra, which we previously discussed, is one example.

The results of Reconstruction, however, calmly point out that behind these individual successes, there is still significant variability in AI's scientific reasoning capabilities depending on the type of task. This result suggests that the ability to construct proofs within established mathematical frameworks is fundamentally different from the ability to conceive entirely new research directions from scratch.

What Researchers Should Consider

The most important lesson this benchmark teaches is that when evaluating AI's contributions to scientific research, it's crucial to continuously verify, with a robust design, "how generalizable its capabilities are," rather than focusing solely on individual, spectacular successes.

In the future, benchmarks targeting "truly new problems that cannot be included in training data" of this type may become more widespread as a standard method for evaluating the scientific reasoning capabilities of AI. We will continue to closely monitor how much the 42% result achieved with the multi-agent method can be improved with further methodological refinements, and whether that growth curve reflects a true "improvement in reasoning ability" or whether it hits another limit.

ベンチマークAI/ML科学的推論研究論文

Monthly commit count doubles—GitHub's 8-hour outage reveals an "unexpected side effect" of the AI ​​coding boom.

This article provides a technical analysis of GitHub's post-incident review (published August 20th) following the 7-hour and 47-minute major outage on August 17th. It explains the auto-scaling design that overlooked the processing limit of the Istio sidecar, the secondary damage caused by a VS Code retry bug that amplified traffic tenfold, the actual growth of monthly commits doubling from 1.4 billion to 2.9 billion in December 2025, and the upward revision of capacity planning from 10 to 30 times.

Monthly Commits Double—GitHub's 8-Hour Outage Reveals an Unexpected Side Effect of the AI ​​Coding Boom

On August 17th, GitHub experienced a massive outage lasting 7 hours and 47 minutes. Almost every service developers rely on daily, including github.com itself, authentication, GitHub Actions, APIs, pull requests, issues, and even Copilot, was affected. The post-incident report, released on August 20th, highlighted, with numbers, not just a temporary problem, but the immense load the surge in AI coding agents is placing on the infrastructure.

The Technical Chain of the Outage

First, let's clarify the technical sequence of events. The outage occurred at 13:28 UTC at GitHub's Central US data center. Traffic reached record highs, causing the Istio sidecar proxy (a widely used microservices architecture technology for managing inter-service communication) to exceed its maximum number of concurrent connections.

The problem here lay in the design of autoscaling (a mechanism that automatically increases or decreases resources based on load). The system's auto-scaling policy was designed based on the main service's processing capacity and did not consider the sidecar's own processing limits. This blind spot meant that even when the sidecar became a bottleneck, the system didn't automatically scale up, causing the capacity shortage to cascade to HAProxy (load balancing software) and authentication paths.

Secondary Damage: A Chain of Retryes Amplifies the Problem Tenfold

Further exacerbated by secondary problems that arose during the recovery process. While partially recovering traffic by diverting some traffic from the central US to the eastern US (Northern Virginia) and investigating the root cause, a delay in response to a single internal endpoint triggered a "potential retry bug" in VS Code.

This bug caused the system to excessively resend failed requests, resulting in a roughly tenfold increase in traffic volume. This amplified wave of retries added further load to a system already in the recovery phase. GitHub explained that they could not fully restore traffic until this behavior could be safely mitigated.

The Nature of the Cause: "Not a Code Change"

GitHub itself clearly emphasizes in this blog post that this outage was not caused by the deployment of new code or configuration changes. The root cause of this outage was pure "insufficient capacity."

This is also the second major incident in August (an Actions outage occurred on August 6th). GitHub frankly admits, "The core of both incidents was a capacity failure. We were unable to scale critical components before demand exceeded processing capacity."

The Reality of Growth: From 1.4 Billion to 2.9 Billion Monthly Commits

The most impressive aspect of this post-incident review was the specific growth rate figures revealed by GitHub. According to the company, monthly commits will double from 1.4 billion to 2.9 billion from December 2025 onwards.

This rapid growth is driven by the rapid adoption of agent-based development workflows. The load generated by pull requests no longer extends beyond simple code reviews, but now spans a wide range of system components, including Git storage, Actions, search functionality, APIs, and background jobs. The fact that AI coding agents are generating commits and pull requests far more frequently than human developers is seen as a contributing factor to this explosive increase in load.

Upward Revision from "10x Capacity Plan" to "30x"

GitHub's CTO also commented on how the company has addressed this challenge. While a plan to increase capacity tenfold had been underway since October 2025, by February 2026, the company had determined that it needed to plan to handle 30 times its current scale in the future.

In other words, GitHub itself had already begun a massive capacity expansion in anticipation of the surge in AI-driven development workflows. However, the pace of that expansion was slightly behind the actual increase in demand, which is the reality behind this outage.

Lessons for Engineers

There are at least two lessons to be learned from this outage. The first point is a subtle but crucial principle: the design of autoscaling in a microservices architecture must be based on the "lowest-performing component" within the entire system. Judging scaling based solely on the main service's processing power carries the risk of easily overlooked bottlenecks, such as the sidecar proxy in this case, causing the entire system to malfunction.

The second point is that retry behavior during failure recovery can unintentionally trigger a "self-DDoS" type phenomenon, exacerbating damage. The retry logic embedded in the client-side implementation (VS Code in this case), combined with a temporary server-side malfunction, can generate so much traffic that it hinders the recovery process itself.

For engineers building workflows utilizing AI agents on their own infrastructure, this GitHub case is a concrete and insightful example demonstrating the significant load that "AI-driven development acceleration" can place on the underlying infrastructure. GitHub aims to complete its plan to migrate all production traffic from its own data centers by 2026.

GitHub障害AIコーディングインフラクラウドインフラ

"5,000 units was always the upper limit"—Tesla itself admits the reality of its 8,000 robotaxi plan in Las Vegas.

On August 20, the Nevada State Transportation Authority unanimously approved Tesla, Uber, and Waymo to operate paid robotaxi services in Clark County, allowing for a maximum deployment of 8,000 vehicles (5,000 Tesla, 1,000 Waymo, and 1,000 Uber). This article skeptically examines the statements made by Tesla's Cybercab chief engineer, who admitted that "5,000 is the upper limit, and around 2,500 is more realistic," as well as the backlash from the taxi industry and the timing coinciding with the Cybercab unveiling event on September 3.

"5,000 Vehicles Was Always the Upper Limit"—Tesla's Own Admittance Reveals the Reality of its 8,000-Vehicle Robotaxi Plan in Las Vegas

On August 20th, the Nevada State Department of Transportation unanimously approved Tesla, Uber, and Waymo's licenses to operate paid robotaxi services in Clark County (the area including Las Vegas). Together, the three companies are expected to deploy up to 8,000 autonomous vehicles over the next 12 months. While this is being reported as one of the largest regulatory approvals in US autonomous driving history, what's noteworthy from a software perspective is the significant gap between the approved "upper limit" and the "realistic outlook" presented by the companies themselves.

The "Asymmetry of Allocation" Revealed by the Breakdown

First, let's examine the breakdown of the authorization. Tesla is allocated a maximum of 5,000 vehicles, Waymo a maximum of 1,000, and Uber a maximum of 1,000 (via Hyundai's Motional and Amazon's Zoox). Zoox has already obtained operating permits for an additional 100 vehicles, bringing the total to approximately 8,000 vehicles in the region.

The exceptionally large allocation to Tesla is interesting in itself. Last year, the company only received a provisional permit for a mere 10 vehicles, limited to the Las Vegas Strip and subject to the condition of having a safety driver on board. In just about a year, the regulatory hurdles have been dramatically lowered to an unconditional permit for 5,000 vehicles.

Tesla's Honest Opinion: "5,000 Vehicles is Unrealistic"

However, the most noteworthy aspect of this article is the extremely frank comment made by Tesla immediately after the approval announcement. Eric Early, Chief Engineer of Tesla's Cybercab, stated during the review process that "5,000 vehicles was always our 'upper limit'."そして「来年のこの時期までに、5,000台を展開できる状況にあるとは思えない。2,500台程度、あるいはそれよりやや多いくらいまで到達できれば、我々としては非常に満足で、十分な成果だと言えるだろう」と続けている。

これは、規制当局から承認された「最大値」と、企業自身が現実的だと考える「実際の目標」との間に、2倍以上の開きがあることを、企業自身の口から認めた発言だ。ヒューマノイドロボットの分野でもたびたび見られる「発表された上限」と「検証可能な実績」のギャップが、自動運転の世界でも、全く同じパターンで繰り返されていることになる。

タクシー業界からの反発という、もう一つの現実

今回の規制承認は、地元のタクシー業界やギグエコノミーのドライバー団体から、強い反発も招いている。Livery Operators Associationや地元タクシー会社の代表者は、この承認が「行き過ぎている」「性急すぎる」と主張し、市場の飽和や道路混雑への懸念を訴えた。

In response to this opposition, Uber has reportedly proposed a "hybrid approach" of gradual integration, arguing that it is preferable for incorporating autonomous vehicles into cities while meeting peak demand. Whether the full 8,000-vehicle limit will be reached all at once, or whether each company will start gradually on a more modest scale, will need to be judged based on the actual pace of operation in the future.

Tesla Cybercab to be unveiled in September

Coinciding with this news, Tesla announced that it will hold a launch event for its "Cybercab" in Austin, Texas on September 3rd. This could be a significant milestone for Tesla's robotaxi project, moving from limited pilot programs using modified Model S/X vehicles to a more full-scale commercial deployment using specially designed vehicles.

The close timing of obtaining permits in Nevada and the unveiling of this dedicated vehicle is surely no coincidence. This can be seen as a clear indication of Tesla's intention to expand its robotaxi business from a limited pilot in Texas to a multi-state commercial network.

What a Software Engineer Should Consider

This regulatory approval is undoubtedly a symbolic step for the autonomous driving industry. However, there is a significant gap between the headline figure of "8,000 units" and Tesla's own practical outlook of "around 2,500 units is more realistic."

Regarding Waymo, while it has already launched paid services in San Francisco, Los Angeles, and Phoenix, and it was previously known that Las Vegas was on the "next list," it has not yet been revealed which vehicle models (Zeekr's Ojai, Jaguar I-Pace, or both) will be deployed, or at what pace.

Regulatory "approval" and the actual "number of units in operation" must always be treated as two separate things in this industry. We will be closely watching how many vehicles each company is actually able to deploy on the roads over the next 12 months, and how local opposition develops.

自動運転ロボタクシーTesla規制Nevada
Advertisement300 × 250