Thursday, August 13, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Evening Edition Free of Charge
TODAY'S LEAD STORY

OpenAI Discovers Two Unknown Vulnerabilities in Chrome – New Model Distributed to Defense Teams to "Win Against Attackers"

On August 10th, OpenAI released "GPT-5.6-Cyber," a specialized model for cyber defense professionals, via Daybreak Red. This article explains the two-stage access structure of Daybreak Blue and Red, the dramatic difference from a 1.5% rejection rate to a 95% response rate, the previously unknown vulnerabilities discovered in the Chrome V8 engine (including CVE-2026-15903), collaboration with industry partners such as Accenture and CrowdStrike, and the mandatory hardware key requirement starting in September, while comparing it with the temporary suspension of Astra.

OpenAI Discovers Two Unknown Vulnerabilities in Chrome – New Model Distributed to Defense Teams to "Win Against Attackers"
(Photo: illustrative)

Two Unknown Vulnerabilities in Chrome Discovered – OpenAI Distributes New Model to Defenses to "Win Against Attackers"

On August 10th, OpenAI released "GPT-5.6-Cyber," an AI model specifically designed for cybersecurity defenders. This announcement comes just days after the company temporarily suspended development of its "Astra" model due to concerns that its cyber capabilities might have reached a "Critical" risk level. This seemingly contradictory two-pronged approach—strengthening vigilance against attack capabilities while actively distributing tools to defenders—is worth examining.

The Defense Professional's Dilemma: Being "Too Rejected"

This model was created in response to a pressing need from the security industry. AI companies like OpenAI have implemented system-level safety measures (guardrails) for cybersecurity-related requests to prevent misuse of their models. However, these safety measures had the side effect of indiscriminately blocking even legitimate defensive work.

Work that appears "aggressive" by its nature, such as penetration testing (a method of intentionally attempting to infiltrate a system with permission to identify vulnerabilities), is highly likely to be rejected by conventional models, even if it is a legitimate defensive activity. This release can be seen as an attempt to directly address this "dual-use" dilemma (the property of the same technology being usable for both good and evil purposes, such as civilian and military, or defense and offense).

Two-tiered access: "Blue" and "Red"

OpenAI has reorganized its "Daybreak" program into two access tiers. "Daybreak Blue" provides approved users with access to a version of the GPT-5.6 Sol frontier model with system-level cyber-related guardrails removed. It is intended for relatively common defensive tasks such as vulnerability discovery, malware analysis, incident response, and patch verification.

On the other hand, "Daybreak Red" provides access to the newly established "GPT-5.6-Cyber" model itself to users who have undergone a more rigorous screening process. This version handles more advanced and high-risk dual-use tasks, such as discovering zero-day vulnerabilities and developing and verifying exploit chains (attack procedures that chain multiple vulnerabilities to ultimately lead to a breach).

Shift from a 1.5% Rejection Rate to 95%

OpenAI illustrates this difference in access tiers with concrete figures. In its internal evaluation of the "Advanced Cybersecurity Completion Rate," which measures the response rate to advanced scenarios such as exploit chain development, authentication bypass, and privilege escalation, standard GPT-5.6 Sol only fulfilled 1.5% of requests. Even with Daybreak Blue access, this figure remained at 2.0%. However, using GPT-5.6-Cyber ​​via Daybreak Red, this response rate jumped to 95.0%.

This extreme difference simultaneously illustrates how strictly the guardrails are designed in normal operation, and how highly practical models that relax these guardrails can be.

Discovering Unknown Vulnerabilities in Chrome

As a concrete example demonstrating the capabilities of this model, OpenAI has announced that GPT-5.6-Cyber ​​discovered two previously unknown vulnerabilities in Google Chrome's JavaScript engine, "V8." One was registered as CVE-2026-15903 and has already been fixed by Google. This vulnerability involved the V8 optimizing compiler incorrectly skipping safety checks when converting values ​​to integers. Exploitation of this vulnerability could allow for memory read/write access, potentially leading to an escape from Chrome's sandbox (a securely isolated execution environment), making it a highly serious issue. The other vulnerability is currently under the collaborative disclosure process (a practice where discoverers and vendors cooperate to keep information confidential until a fix is ​​complete).

OpenAI further explains that it has used this model to discover five critical vulnerabilities in mobile operating systems, three in a database, and over 400 potential privilege escalation vulnerabilities in an OS kernel.

Collaboration with Industry Partners

The Daybreak Cyber ​​Partner Program includes renowned security companies such as Accenture, CrowdStrike, Cisco, IBM, and Palo Alto Networks. Harpreet Sidhu, Global Cybersecurity Lead at Accenture, commented, "Security teams are now under pressure not only to quickly find vulnerabilities, but also to quickly fix them." The challenge going forward will likely be how much the speed from discovery to fix can be reduced through the power of AI.

Furthermore, from September 1st, the use of hardware security keys will be mandatory for individual Daybreak accounts. This reflects the stance that access to a powerful tool must be accompanied by correspondingly enhanced identity verification.

What Engineers Should Consider

This release indicates not a simple question of "how to limit AI's cyber capabilities," but rather a more complex control design: "to whom and to what extent should restrictions be relaxed?" The idea of ​​providing defenders with equivalent or superior tools before attackers can launch automated AI attacks can be seen as an attempt to rectify, to some extent, the "asymmetry between attack and defense" in the world of cybersecurity.

On the other hand, it remains to be seen how strictly access to highly capable models like GPT-5.6-Cyber ​​will continue to be controlled. This series of actions—the temporary suspension of Astra and the provision of GPT-5.6-Cyber ​​to defenders—suggests that OpenAI is simultaneously pursuing both "suppression of attack capabilities" and "enhancement of defensive capabilities."

OpenAIサイバーセキュリティAI/ML脆弱性セキュリティツール

Unitree finally gets a stock price – a miscalculation behind its $9.04 billion valuation: a "changing of the guard"

Unitree Robotics' stock price was set at 150.8 yuan on the Shanghai STAR market, setting a valuation of approximately $9.04 billion for its IPO. The high level of interest, with individual investors having an allocation rate of only 0.018%, the fact that Agibot holds the top spot with a 44% share in global humanoid robot shipments of 19,100 units in the first half of 2026 (a 272% increase year-on-year), the reality that Chinese manufacturers account for over 97% of shipments, and the inflow of $13.86 billion into the embodied intelligence sector, are all examined as indicators of the shifting power dynamics behind this milestone of listing.

Unitree Finally Sets a Share Price – A Miscalculation: A "Shift in the Top" Behind a $9.04 Billion Valuation

Unitree Robotics, which has been preparing for a listing on the Shanghai STAR Market, has set its share price at 150.8 yuan (approximately $22.34) per share. This brings the company's valuation to approximately 61 billion yuan (approximately $9.04 billion), and it aims to raise 610 million yuan. While we covered the listing preparations in our previous article, this time we will look at the new stage of price setting and the shift in the industry's power dynamics that has become apparent simultaneously.

The "Unusual Overheating" Indicated by the Application Ratio

What stands out in this IPO is the allocation rate for individual investors. According to documents filed with the exchange, the final allocation rate for retail investors was approximately 0.018%. This means that only about 1 in 5,500 individual investors who applied received shares. The Chinese IPO market is inherently prone to overheating, but the enthusiasm for investment in humanoid robot companies is clearly evident from these figures.

The company states that it will use the funds raised for robot software and hardware development, new product launches, and expansion of manufacturing capacity. Interestingly, the company's prospectus itself lists US regulations as a clear risk factor for its overseas expansion. The fact that the FCC regulations at the end of July are casting a shadow over the company just as its humanoid business is projected to become its largest business area in 2025 is nothing short of ironic.

Unitree Falls from "Top Spot"

However, following the news related to this IPO has revealed more essential information. According to research by Smart Analytics Global (SAG), global shipments of humanoid robots in the first half of 2026 are projected to reach approximately 19,100 units, a 272% increase compared to the same period last year. Up to this point, these figures are not surprising as they demonstrate the rapid growth of the industry as a whole. However, what is noteworthy is the fact that it was "Agibot," not Unitree, that took the top spot with a 44% market share.

Until now, when compared to Western companies, Unitree has always been described as "the leading Chinese humanoid robot." However, this data indicates that the competitive structure within China has already shifted. The fact that Unitree had already relinquished its top position in the industry at the very moment it reached the major milestone of its IPO should be an undeniable factor for investors evaluating this IPO.

The Industry Reality: 97% Made in China

Another noteworthy point from the SAG survey is that Chinese manufacturers account for over 97% of shipments. Industrial and commercial applications account for over 70% of total shipments, and SAG predicts that annual shipments will reach nearly 60,000 units, with the industry's total revenue reaching approximately $1.6 billion.

Judging from these figures, it's undeniable that Chinese manufacturers have already established an overwhelming position in terms of "mass production" of humanoid robots. However, evaluating the "quality" of these "mass-produced" robots—how much autonomy they possess that makes them practical, or how long they actually remain operational—is another matter entirely. It would be premature to judge technological superiority solely based on quantitative indicators such as the number of units shipped.

Accelerating Fundraising in China

In parallel with Unitree's IPO, the inflow of funds into other Chinese robotics companies is also gaining momentum. According to ITjuzi data, funding for the "embodied intelligence" sector in China reached 93.5 billion yuan (approximately $13.86 billion) in the first half of 2026 alone, representing a five-fold increase year-on-year through 322 transactions. Funding is also flowing into companies other than Unitree, such as Noin Intelligence, PokeBot, and Discover Robotics.

The Chinese government's "15th Five-Year Plan" places robotics and embodied AI at the heart of its industrial strategy, and this policy support appears to be further accelerating the momentum of fundraising.

What Software Professionals Should Watch

Unitree's IPO is undoubtedly a significant milestone for the humanoid industry. However, the fact that Agibot has now taken over the top spot indicates that the landscape of this industry is changing at a much faster pace than investors and the media anticipate.

It's important to keep in mind that investing in Unitree doesn't simply mean investing in the entire Chinese humanoid market. We should closely monitor the financial data disclosed after the IPO to see the actual pace at which the company is expanding and what position it can maintain amidst the intensifying domestic competition.

UnitreeヒューマノイドIPO中国フィジカルAI

How to create a "yardstick" to measure cyber capabilities—the details of the evaluation methodology released by OpenAI

This article analyzes OpenAI's proprietary evaluation metric, "Advanced Cybersecurity Completion Rate," released in its August 10th GPT-5.6-Cyber ​​announcement, from the perspective of evaluation methodology design. It explains the concept of measuring "whether a response was made without rejection" rather than the accuracy rate, the multi-layered guardrail control shown by the change in response rate from 1.5% for the standard model to 95% for red access, the criteria that differentiated the High and Critical classifications in the Preparedness Framework from Astra, the empirical verification through the discovery of Chrome vulnerabilities, and its significance as a governance method for dual-use research.

How to Create a "Measuring Stick" for Cyber ​​Capabilities: The Details of OpenAI's Released Evaluation Method

On August 10th, when OpenAI announced its cybersecurity-focused model "GPT-5.6-Cyber," it also explained its unique internal evaluation method for measuring the model's cyber capabilities: "Advanced Cybersecurity Completion Rate." As a journalist with a research background, I would like to focus on the evaluation design itself—how to measure the model's capabilities.

The Concept of Using "Rejection Rate" as an Evaluation Metric

General AI model benchmarks often measure "accuracy," which is how accurately a model can perform a specific task. However, the metric used by OpenAI is different. It measures the "completion rate"—how much the model responded (how often it didn't reject) to advanced, high-risk scenarios such as exploit chain development, authentication bypass, and privilege escalation.

This design philosophy stems from a fundamental dilemma inherent in the security mechanisms of frontier AI models. The problem is that guardrails, designed to enhance safety, simultaneously hinder their use for legitimate defensive purposes. Introducing an evaluation axis measuring "how cooperative (or uncooperative) the model is," in addition to the traditional "how accurately the model performs the task," represents an interesting development in dual-use technology evaluation methods.

The Numbers Reveal the Effectiveness of Guardrails

The numbers released by OpenAI clearly illustrate the significance of this evaluation axis. The standard GPT-5.6 Sol responded to only 1.5% of the requests in this evaluation. Even with Daybreak Blue, which has relaxed guardrail requirements, this figure remains at a mere 2.0%. However, using GPT-5.6-Cyber ​​via Daybreak Red, which has undergone more rigorous testing, the response rate jumps to 95.0%.

This dramatic difference demonstrates that safety mechanisms can be designed not as a single on/off switch, but as a multi-layered control system that can be adjusted incrementally according to user confidence. However, a high response rate of 95% also means that the risk of misuse is equally high. It is necessary to evaluate this model with a correct understanding of the trade-off between "high response rate" and "impact on security."

Positioning within OpenAI's Preparedness Framework

The capability evaluation of this model is also related to OpenAI's risk management framework, the "Preparedness Framework." Interestingly, according to Axios's report, GPT-5.6-Cyber ​​remains in the "High" risk category, and does not reach the highest category of "Critical," which Astra, whose development was temporarily suspended a few days ago, may have reached.

Despite both being models dealing with cybersecurity capabilities, one is being offered to defenders for public release, while the other's development has been temporarily suspended—a stark contrast. This difference lies precisely in the risk classification criteria defined by the Preparedness Framework. The fate of both models hinged on whether they could autonomously discover and exploit zero-day vulnerabilities in robust real-world systems without human intervention.

Empirical Verification: "Discovering Chrome Vulnerabilities"

As evidence that this model's evaluation goes beyond mere theoretical benchmarking, OpenAI cites the fact that GPT-5.6-Cyber ​​actually discovered two previously unknown vulnerabilities in Google Chrome's V8 engine. One of these was registered as CVE-2026-15903 and has been fixed by Google. This vulnerability involved the V8 optimizing compiler mistakenly skipping safety checks during integer conversions, a serious issue that could lead to memory corruption and sandbox escape.

Such real-world results serve as compelling evidence connecting abstract benchmark scores to concrete technical achievements. However, while these results demonstrate the model's capabilities, they also simultaneously prove the high risk of exploitation, a point that the research community must calmly acknowledge.

The Challenge of Governance in Dual-Use Research

OpenAI's recent initiative offers a practical solution to the "dual-use" problem in AI research. Instead of completely disclosing or completely sealing off technologies with the same capabilities, it designs a system where access is gradually opened up based on the user's level of identity verification.

This type of access control model is similar to the management methods for "dual-use research of concern" that have been debated for many years in the fields of biology and chemistry. The policy of mandating identity verification using hardware security keys from September can also be understood in this context.

What Researchers Should Consider

OpenAI's combination of "rejection rate as an evaluation axis" and "gradual access control" indicates that the security evaluation of frontier AI models is evolving from simple performance measurement to a more complex framework that includes social operational design.

Going forward, it will be interesting to see to what extent this type of "completion rate"-based evaluation method and hierarchical access control become standard practice when other AI development companies deploy similar dual-use models. Furthermore, independent third-party verification of how much these evaluation metrics actually contribute to the safe operation of the models will likely become an important research topic in the future.

OpenAIAI安全性評価手法デュアルユースAI/ML
Advertisement300 × 250