Can AI Recreate Ideas from "Unpublished Papers"? – The Real Barrier of Scientific Reasoning Revealed by a Figure of Only 3-15%
In August, a new benchmark called "Reconstruction" was published on arXiv. It's a unique evaluation method that shows Frontier's large-scale language model only the bibliography of a research paper and asks it to reconstruct what hypotheses the paper actually presented. The results were surprisingly low, with even the best-performing models scoring only 3-15%. As a journalist with a background in AI research, I want to carefully examine what these results mean.
A Clever Design That Guarantees "Not Remembering the Answers"
The technical strength of this benchmark lies in its direct confrontation with a deep-seated problem in AI evaluation: "benchmark contamination." Benchmark contamination refers to the phenomenon where the questions and answers intended for evaluation become embedded in the model's training data, causing the model to essentially "memorize the answers" before taking the test. A saturation survey conducted in 2026 on 60 LLM benchmarks indicated that nearly half of them were potentially already "saturated" by this type of contamination.
Reconstruction employs a design that structurally avoids this problem. If the evaluated paper was published after the model's training data cutoff date, the possibility of its content being included in the model's training data is, in principle, zero. In other words, the model must not "remember" the content of the paper, but purely "infer" what the paper actually argued based solely on the clue of the bibliography.
The Harsh Reality: "3-15% on its own"
The result measured under this clever design is the 3-15% score mentioned at the beginning. This figure indicates that the frontier model's ability to accurately reconstruct the research idea itself from indirect clues such as bibliographies is currently extremely limited.
The research team argues that this result reflects the true capability ceiling the model faces, rather than being an apparent low due to poor prompt design or some flaw in the evaluation method. The figure is compelling precisely because the measurements were taken in a contamination-free environment.
Multi-agent approach improves to 42%—still not quite half
Interestingly, the research doesn't end there. The research team also evaluated the same task using a "multi-agent approach" combining multiple AI models. Specifically, they used a pipeline combining mutual review by multiple models with a "Swiss tournament" (a highly fair tournament format where the next opponent is determined by wins and losses) to select the top four candidate options. They did not use external web searches, relying solely on the provided reference information for inference.
This collaborative approach with multiple models raised the accuracy to 42%. This is a significant improvement compared to the best score of a single model. However, it still remains below half. This suggests that even combining multiple AIs still doesn't reach the level of "expert scientific intuition." ## The Difficulty of Taking These Results "At Face Value"
On the other hand, caution is needed when interpreting these research results. Whether the 15% figure, which achieved the top score, truly constitutes "inference" in a philosophical sense remains debatable. It's impossible to determine from this benchmark score alone whether it simply outputs statistically "plausible" hypotheses from combinations of references, or whether it involves some meaningful reasoning process.
However, the "practical ceiling" demonstrated in this paper—the fact that current models have clear limitations in their ability to reconstruct truly new research ideas from reference clues—can be empirically supported by the robust design of contamination resistance.
The Irony of Timing: Precisely During a Series of High-Profile Announcements
The timing of this paper's publication is also highly suggestive. In August 2026, several Frontier Institutes announced a series of spectacular scientific achievements, such as AI solving unsolved mathematical problems and discovering vulnerabilities in cryptographic algorithms. The mathematical proof by OpenAI's Astra, which we previously discussed, is one example.
The results of Reconstruction, however, calmly point out that behind these individual successes, there is still significant variability in AI's scientific reasoning capabilities depending on the type of task. This result suggests that the ability to construct proofs within established mathematical frameworks is fundamentally different from the ability to conceive entirely new research directions from scratch.
What Researchers Should Consider
The most important lesson this benchmark teaches is that when evaluating AI's contributions to scientific research, it's crucial to continuously verify, with a robust design, "how generalizable its capabilities are," rather than focusing solely on individual, spectacular successes.
In the future, benchmarks targeting "truly new problems that cannot be included in training data" of this type may become more widespread as a standard method for evaluating the scientific reasoning capabilities of AI. We will continue to closely monitor how much the 42% result achieved with the multi-agent method can be improved with further methodological refinements, and whether that growth curve reflects a true "improvement in reasoning ability" or whether it hits another limit.