Tuesday, August 4, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Morning Edition Free of Charge
TODAY'S LEAD STORY

OpenAI's unreleased model "Astra" solved 10 unsolved mathematical problems for $2,000.

On August 1st, OpenAI announced that its unreleased next-generation flagship model, "Astra," had solved 10 problems in the fields of mathematics and theoretical computer science that had remained unsolved for over 10 years, and had released them along with machine-verifiable proofs in Lean 4 format. This article explains the differences from the erroneous announcement regarding the Erdős conjecture in October 2025, the solution of the 27-year-old unsolved problem of "non-Sophic groups" in group theory, the computational cost of approximately $2000, and how OpenAI is addressing the risks pointed out in the Leiden Declaration.

OpenAI's unreleased model "Astra" solved 10 unsolved mathematical problems for $2,000.
(Photo: illustrative)

OpenAI's Unreleased Model "Astra" Solves 10 Unsolved Mathematical Problems for $2000

On August 1st, OpenAI made an unusual announcement. Their next flagship model, "Astra," which is not yet publicly available, solved 10 problems in the fields of mathematics and theoretical computer science that had remained unsolved for over 10 years, and released all of them with "machine-verifiable certificates." For engineers, what makes this news interesting is not the results themselves, but "why this time it can be believed."

The Deep-Seated Distrust of "Another AI Solves Math Problem"

Many engineers must have felt a sense of déjà vu upon hearing this news. In fact, in October 2025, an OpenAI executive at the time announced that "GPT-5 had solved 10 unsolved problems of the Erdős conjecture." However, this announcement crumbled within days. Mathematician Thomas Bloom's verification revealed that the model had merely "searched" for solutions from existing literature and had not constructed its own proofs. What makes Astra's announcement different is that Bloom himself has described it as "important news." The mechanism that convinced him and the entire mathematics community is the core of this announcement.

Providing proofs in a "verifiable" format, not just a "readable" one

The key technical point is that all ten proofs were formalized using Lean 4, a "theorem proving assistance system" (a tool that allows mathematical proofs to be written in a format that a computer can verify line by line, like a program).

Normally, mathematical proofs generated by AI can only be judged by human mathematicians as "likely correct." However, with Lean-format proofs, readers can mechanically verify whether the proof is truly correct simply by running the code on their own laptops. The repository released by OpenAI reports zero occurrences of the reserved word "sorry," which indicates incomplete parts of the proof—meaning that all logical steps are complete and without omissions.

This represents a significant practical shift from "believing" in AI output to "verifying" it.

The Highlight: A Group Theory Problem Unsolved for 27 Years

The 10 results span a wide range of fields, including group theory, von Neumann algebras, higher-dimensional geometry, quantum computation, lattice cryptography, and extreme value combinatorial theory. The highlight is the first concrete example of a group construction called a "non-Sophic group."

This was a central unsolved problem in group theory, which no one had been able to prove or disprove for 27 years since mathematician Mikhail Gromov proposed the concept of "Sophicity" in 1999. This problem has now been finally resolved.

The Significance of a Computational Cost of Only $2000

According to OpenAI's research leader, Sebastian Bubek, the computational cost to derive all 10 solutions was approximately $2000 in terms of the company's API fees. This is a symbolic figure. A problem that has baffled human intellect for decades has been solved for the cost of a few hundred cups of coffee.

However, there are points to keep in mind before blindly accepting this figure. This is merely the "cost of the tokens used to derive the solution, based on their list price," and does not include the enormous research and development costs invested in training the model.

Mixed Welcome and Cautious Expectations from the Mathematical Community

This result is not simply being consumed as self-proclaimed by AI companies. In June 2026, the "Leiden Manifesto," a document endorsed by the International Mathematical Union and signed by prominent mathematicians such as Terence Tao, Peter Scholz, and Scott Aaronson, was published. This manifesto lists five risks in the relationship between AI and mathematics: low reliability of results, lack of sources, reliance on closed commercial systems, exaggerated claims, and loss of scientific independence.

Astra's announcement is designed to directly address at least three of these five risks. The proof is verifiable in Lean format, and a 249-page technical document and a 62-page explanatory document detailing how the model constructed its reasoning have been released. The stance isn't "just show us the results and believe it," but rather "we'll give you the verification process, so feel free to question it."

What Engineers Should Consider

The essence of this news isn't that "AI can now solve math problems." It's that "a system is beginning to take shape that allows third parties to independently and mechanically verify the results produced by AI."

When integrating the output of generative AI into practical work, many engineers face the hurdle of "how to verify whether this output is truly correct." While this case is a success in the highly verifiable field of mathematics, the idea of ​​"how to provide verifiable backing for AI output" is likely to be applied to more familiar development environments, such as code generation and design document reviews.

OpenAI has also announced that it will provide 100,000 academic researchers with free access to its frontier models until 2027. While this could be seen as a move to lock research infrastructure into its own platform, if the culture of "verifiable proof" becomes the standard way in which AI and mathematical research interact, then this is a welcome direction in itself.

OpenAIAstraAI/ML数学Lean形式検証

Giving up cars to build robots—a look inside BYD's "first prototype" to be launched in August.

Chinese EV giant BYD announced it will unveil its first humanoid robot at its "Di Space" experience facility in August. This article offers a skeptical look at the commonalities in the technology stacks behind the successive entry of Chinese automakers such as Tesla, Xiaomi, and XPENG into the humanoid business, the difficulty in verifying the 2026 shipment forecast of 100,000 units, and the current situation of the Tesla Optimus, whose production start date at the Fremont factory has been pushed back to "the second half of this year" as a point of comparison.

Giving Up Cars to Building Robots: Examining the Inside of BYD's First Robot to Launch in August

Chinese EV giant BYD has officially announced that it will unveil its first humanoid robot in August. The announcement will take place at the company's experience facility, "Di Space," where visitors will be able to interact with the robot. While it's tempting to dismiss this as "another EV manufacturer entering the robot market," a closer look reveals a movement that cannot be ignored within the industry structure.

Why Automakers are Rapidly Entering the Humanoid Market

Following BYD's lead, Chinese automakers such as Tesla, Xiaomi, XPENG, FAW, Chery, Dongfeng, Changan, and GAC are all venturing into the humanoid business. This is no coincidence. Self-driving cars and humanoid robots share a significant overlap in their technological stacks: "perceiving and making decisions using sensors, and physically moving with motors and actuators."

In fact, Tier 1 suppliers of automotive parts—Joyson Electronics, RoboSense, Hesai, and Black Sesame—are also showing signs of adapting automotive-grade technology for robotics. Joyson Electronics showcased a full stack of products at the 2026 World Artificial Intelligence Conference (WAIC), including its dexterous hand "Lingxi," solid-liquid hybrid batteries, third-generation AI head modules, and electronic skin. Therefore, BYD's entry should be seen as part of a trend to liberate the automotive industry's "recognition, judgment, and driving" know-how from the confines of the automobile.

Phase Shift from "Mass Production Verification" to "Commercial Pilot"

Industry forecasts suggest the humanoid market is rapidly shifting from the technology verification phase to the commercial pilot phase, with global shipments potentially reaching 100,000 units by 2026. If this is true, it would represent a crucial turning point toward large-scale deployment. What I want to emphasize here is that figures like "tens of thousands of units" are often industry forecasts or self-reported by companies, and in most cases, they haven't been verified by an independent third party. Especially with announcements related to robots originating from China, the numbers tend to circulate without much backing from actual operational performance or mass production systems. Regarding BYD's announcement, it's necessary to distinguish between the fact that it will be unveiled in August and how far it will actually be put into practical use and mass production.

Comparison: How far has Tesla's "Optimus" progressed?

A comparable example of a humanoid robot originating from the automotive industry is Tesla's Optimus. In its January 2026 earnings announcement, Tesla announced the end of Model S/X production at its Fremont plant and its conversion to the Optimus production line, initially guiding it to "begin low-volume production from the end of July to August."

However, in the shareholder update on July 22nd, this specific timeframe was replaced with the more vague phrase, "sometime later this year." The production report as of July 2nd also makes no mention of Optimus production figures. Musk himself has repeatedly warned that the initial production pace will be "very slow," and given that it's an entirely new production line using over 10,000 custom parts, this delay is not surprising.

While competitors such as Boston Dynamics' Atlas, Figure's Figure 03, and Agility Robotics' Digit have already accumulated verifiable operational records at sites like Hyundai, BMW, and GXO (Digit boasts over 65,000 hours of operation at nine commercial facilities), Tesla has yet to release any official figures for "operational units." When looking at the humanoid industry, it's best to always treat "announced production start dates" and "independently verifiable operational records" as separate things.

What "Functional" Means: A Software Engineer's Perspective

BYD has stated that this time, they will be showcasing a "functional robot," not just a concept exhibit, and creating an interactive exhibit where visitors can actually touch it. Personally, this is the point I'm most interested in. Is it just a demo-level functionality like "walking" or "waving arms," ​​or does it possess a level of autonomy closer to practical use, such as recognizing and grasping objects, or performing tasks according to simple instructions? This difference cannot be judged solely from the presentation.

In the past, there have been many cases in this industry where "robots that moved smoothly in demos were actually remotely controlled." When evaluating BYD's presentation, it will be necessary to carefully determine from the video and on-site reports what was done "autonomously" and what was "operator-assisted" during the August demonstration.

What Engineers Should Look For

The rush of humanoid robots from the automotive industry is a rational trend in terms of the horizontal deployment of the "recognition, judgment, and driving" technology stack. BYD's entry can be understood in this context.

On the other hand, we should remain as wary as ever of the "numbers taking on a life of their own" that often accompanies announcements related to robotics originating from China. What we should really be looking at in BYD's August announcement is not the excitement at the exhibition hall, but the mundane but essential question of "how much of it is autonomous operation, and how much is human assistance?" I would like to re-evaluate it once demonstration videos and verification reports from third parties become available.

BYDヒューマノイド中国TeslaOptimusフィジカルAI

Is an AI with a "high safety score" truly safe? — A paper that examined the benchmark itself raises this question.

The paper "Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks," published on arXiv on July 30, validates the major safety benchmarks R-Judge, InjecAgent, AgentHarm, and AgentDojo using a panel of 9 organizations and up to 22 models. It discusses flaws in the metrics that can produce high scores even with simple judgment methods, ranking inconsistencies between benchmarks, the negative correlation between capability and misalignment safety (ρ=-0.44) and how it changes with panel expansion, and the convergent validity of AgentHarm and jailbreak resistance.

Is an AI with a High Safety Score Truly Safe? – A Question Raised by a Paper That Examines the Benchmarks Themselves

Numerous benchmarks have been proposed to measure the safety of AI agents. R-Judge, InjecAgent, AgentHarm, AgentDojo – many researchers and engineers have likely heard of these. However, the paper "Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks," posted to arXiv on July 30th, raises a fundamental question: "Is it really appropriate to directly cite the scores of these benchmarks as an indicator of 'AI agent safety'?"

Questioning the Benchmark as a Measuring Stick

What makes this paper unique is that it examines not the AI ​​model itself, but the benchmark—the "measuring stick" used to measure AI models. The research team ran four major safety benchmarks using official implementations and scorers on up to 22 models across nine development organizations (OpenAI, Meta, Qwen, Mistral, DeepSeek, Amazon, Anthropic, Google, and Cohere). They also independently measured scores on MMLU and GPQA, two capability benchmarks, to create a "comprehensive capability index."

In other words, this research re-examined claims such as "this model scored highly on safety benchmarks," rather than blindly accepting them.

The problem of achieving high scores simply by "always judging something as dangerous"

The first point raised is a flaw in the evaluation metric itself. The R-Judge benchmark judges AI behavior records (traces) as binary (safe or dangerous) and evaluates them using an index called the F1 score. However, according to the analysis in this paper, even with a policy that essentially makes no judgments—"always judging something as 'dangerous'"—the F1 score reaches 0.690. This score is higher than that of 5 out of 21 models that actually possess discriminative ability.

This illustrates a typical pitfall in evaluation methods: depending on the design of the metrics used for evaluation, a "do-nothing" judgment method can appear "excellent."

Discrepancies in Benchmark Rankings

Even more problematic is the fact that multiple broad-coverage benchmarks assign different rankings to the same 18 models. The paper identifies that this discrepancy is due to "artifacts (apparent statistical patterns) specific to small panels."

Specifically, the correlation coefficient between R-Judge's specificity (the ability to correctly judge safe behavior as safe) and AgentHarm's safety score changes from -0.64 in a small sample of 7 models to almost zero at +0.02 when expanded to 18 models. Furthermore, when 7 models are randomly selected, a strong correlation, where the absolute value of the correlation coefficient is 0.5 or higher, appears by chance in one-quarter of the cases, even though it shouldn't exist. This study numerically demonstrates that a small number of models being evaluated increases the risk of misinterpreting statistical noise as a "meaningful trend."

The Paradoxical Correlation: "Higher Capability AI is More Dangerous"

The most controversial aspect of this paper is the relationship between capability and misalignment (where AI deviates from its intended behavior). While the overall capability index shows a positive correlation with actual task success rate (ρ=+0.60), it shows a negative correlation with misalignment safety scores (ρ=-0.44, n=21).

When viewed across a panel of 20 paired models, this difference reached a statistically significant level of Δ=-1.00 (95% confidence interval [-1.48, -0.49], p<0.001). This trend remained even after recalculating by excluding organizations one by one, and after performing bootstrap analysis at the organizational level. Interestingly, however, when the study was expanded to 41 models, this negative correlation weakened to -0.16 (confidence interval crosses 0 [-0.54, +0.22]), while the correlation with jailbreak resistance strengthened to +0.34—although neither change was statistically significant, as reported.

In other words, it is premature to conclude with the simplistic modeling of "higher-performing models are more dangerous," as the conclusion can change depending on the scope of the models studied and the type of safety being measured.

The Meaning of "Convergent Validity" Demonstrated by AgentHarm

On the other hand, there is also some positive news. The AgentHarm benchmark showed a strong correlation of ρ = +0.72 with jailbreak safety tests using three different templates, even after statistically removing the influence of performance.

However, the paper also includes a cautious reservation here. The two metrics both measure something of similar nature—"how often an AI follows harmful instructions"—and therefore, this is not evidence that it accurately measures the overall safety of the AI ​​agent. Rather, it merely demonstrates a more limited sense of validity (convergent validity)—that two similar metrics are consistent with each other.

What Researchers Should Consider

The message this paper ultimately conveys is extremely simple: "If you're going to claim safety, you must at least specify which benchmark, which metrics, what was measured, and which model group was targeted."

While this may seem obvious, it's information that many safety announcements and news articles omit. When you see a sentence like, "This model achieved the top score on the safety benchmark," the habit of checking which benchmark and which metrics are being referred to will become increasingly important as you follow the still-developing field of AI safety evaluation.

References: arxiv.org / arxiv.org
AIエージェントAI安全性ベンチマークAI/ML論文統計
Advertisement300 × 250