Thursday, August 20, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Morning Edition Free of Charge
TODAY'S LEAD STORY

AI "accidentally confesses" its weakness—an unprecedented discovery method exposes Copilot's vulnerability.

On August 18th, Microsoft fixed the CoSnitch vulnerability (CVE-2026-24301) in Copilot Personal. This article explains the chain of three vulnerabilities: automated prompt execution via undisclosed URL parameters, data exfiltration from authorized OAuth-integrated applications, and indirect prompt injection via web summarization. It also details the discovery method called "meta-hacking," developed by Varonis Threat Labs, which involves bombarding Copilot with questions to extract its internal structure, and the fact that this is the third such vulnerability this year, following Reprompt and SearchLeak.

AI "accidentally confesses" its weakness—an unprecedented discovery method exposes Copilot's vulnerability.
(Photo: illustrative)

AI "Accidentally Confesses" Its Weakness—An Unprecedented Discovery Method Exposes a Copilot Vulnerability

On August 18th, Microsoft released a patch to fix a serious vulnerability called "CoSnitch." This vulnerability allows attackers to silently steal data from a victim's linked accounts, such as email, cloud storage, and calendar, simply by clicking a single link. While the technical severity is noteworthy, what was most interesting to me as an engineer was how this vulnerability was discovered.

An Attack Chain Linking Three Vulnerabilities

CoSnitch (CVE-2026-24301) is an attack that combines three different vulnerabilities present in Microsoft Copilot Personal. The first is the combination of Copilot's standard query parameter "?q=" with an undocumented, undisclosed parameter, allowing the attacker to automatically execute a prompt upon page loading, without any clicks or confirmations.

The second method involves using these automatically executed prompts to retrieve data from OAuth applications (such as Gmail, Google Drive, and Calendar) that the user has already granted access to Copilot, and then silently sending it to a server prepared by the attacker. It's crucial to note that this is not "unauthorized OAuth access." The victim had already granted Copilot access to these linked applications. CoSnitch broke the implicit assumption that "operations on linked applications only occur at the user's intentional request."

The third method is an indirect prompt injection (an attack technique that injects malicious instructions into normal content) that exploits the web summarization feature. The attacker prepares a seemingly harmless webpage by embedding malicious instructions in HTML comments, metadata, or visually hidden text. When a victim requested a summary of this page from Copilot, Copilot processed not only the displayed text but also hidden attacker instructions, resulting in malicious instructions being written to the user's "persistent memory."

Discovered New Technique: "Meta-Hacking"

The technique used by security company Varonis Threat Labs, which discovered this vulnerability, is attracting attention in the industry. The company calls this "meta-hacking." While this type of vulnerability is usually discovered through reverse engineering of the code, Varonis researchers took an approach of repeatedly asking Copilot "why automated execution is impossible."

Each time Copilot refused to answer the question, the explanation for the refusal included information about the system's technical workings. The researchers restructured each refusal into further questions, gradually refining the accuracy of the information that could be extracted from Copilot. In the process, Copilot itself, unsolicited, disclosed undisclosed URL parameters that would be used in the attack, midway through its explanation of the refusal. Varonis described this series of interactions as "a combination of highly sophisticated social engineering against LLMs, various jailbreaks, and prompt injection attacks."

AI Assistants Create Their Own "Map of Attack Surfaces"

This meta-hacking technique demonstrates the risk that the very act of an AI assistant attempting to explain its own security can unintentionally provide attackers with useful information. Refusal messages like "This cannot be done" are often accompanied by technical reasons for their inability. As these explanations accumulate, attackers can reconstruct the system's internal structure through dialogue alone, without ever examining the code.

This represents a new type of risk, distinct from conventional software security practices. While traditional bug detection involved analyzing static code and communication packets, in the case of AI assistants, clues to vulnerabilities can be extracted by "conversing" with the system itself.

A Third Vulnerability in Copilot

For Varonis, CoSnitch marks the third Copilot-related vulnerability discovery this year. In the past, Varonis has reported "Reprompt," which bypasses security mechanisms by repeatedly asking the same question, and "SearchLeak," which turns Microsoft 365 Copilot Enterprise into a covert data exfiltration route. All three vulnerabilities share an extremely simple starting point for attacks: "single click on a seemingly normal link."

Varonis reported this issue to Microsoft in December 2025, and it took approximately eight months for the actual patch to be released. Fortunately, no evidence of this vulnerability being exploited before the patch was released has been found.

What Engineers Should Consider

The lesson this news teaches is that careful auditing is necessary when designing AI assistants to have access to a wide range of data, such as email, calendars, and cloud storage. Varonis recommends that corporate security teams regularly audit third-party applications connected to Copilot, treat AI assistants as "authorized insiders requiring the same level of access oversight as human employees," and establish monitoring systems to detect anomalous data access via AI assistants. When integrating an AI assistant into your company's development workflow, it's worth taking stock of how much data access you're granting, and under what conditions, before considering its "convenience."

MicrosoftCopilotサイバーセキュリティ脆弱性AIエージェント

"Monitoring costs account for 20% of inference computation"—OpenAI reveals the design philosophy behind its multi-stage monitoring system.

This analysis examines OpenAI's technical blog post, "Pacing model development in an era of cyber-critical capabilities," published on August 18th. It explains, from a technical architecture perspective, the escalation structure from an activation classifier that checks at the token level to an automated investigator, the specific operational criteria of 30-minute alerts + 30-minute judgments, the cost estimate of 20% of the monitored computational load and the policy of not passing this cost on to customers, the Preparedness Framework rewriting plan formulated in 2023, and the difference in approach with Anthropic.

"Monitoring Costs 20% of Inference Computation"—OpenAI Reveals the Design Philosophy of its Multi-Stage Monitoring System

On August 18th, OpenAI published a detailed technical blog titled "Pacing Model Development in an Era of Cyber-Critical Capabilities." This is the first concrete explanation of the technical design of the monitoring system that underpinned the decision to temporarily suspend development of the Astra model, which we previously discussed. As a journalist with a background in AI research, I would like to carefully examine the details of this monitoring architecture.

A Multi-Stage Monitoring Pipeline Starting with "Activation Classifiers"

OpenAI's new monitoring system is designed not as a single checking mechanism, but as a multi-layered pipeline with multiple stages. The first stage is a detector called "activation classifiers." This mechanism checks the internal activity state of the model for each word (each sampled token) generated by the model.

If any concerning signs are detected during this initial check, the signal is escalated to "automated investigators" with more advanced computing resources. These automated investigators examine the tool operations the model attempted to perform, the available reasoning processes, and the entire sequence of actions to conduct a deeper analysis for unauthorized access, data exfiltration, destructive behavior, or attempts to circumvent security mechanisms.

The "30-Minute Rule": Specific Operational Standards

The practical operational rules for this monitoring system are also publicly available, along with specific figures. OpenAI aims to issue an alert within 30 minutes of detecting concerning signs. If the alert is determined to be "highly likely to be a breach of a critical security boundary," the safety, security, and research teams are simultaneously paged (requested to act as paging agents).

These teams are then given another 30 minutes. If they cannot determine within this time that the alert is a false positive, the team is required to temporarily suspend the relevant training run or evaluation work. This decision-making process, with its clear time limit of "30 minutes + 30 minutes," can be seen as a concrete institutional design to prevent concerns about the security of AI models from being left unresolved.

Cost: "20% of monitored computational resources"

Operating this monitoring system naturally requires computing resources. OpenAI estimates the overhead of this monitoring to be "approximately 20% of the monitored inference computational resources." However, it is added that this cost varies significantly depending on the type of workload.

This monitoring is mandatory for all reinforcement learning (RL) training and evaluation work using the tool for models with capabilities equivalent to or greater than Sol. Furthermore, after it was determined on August 7th that the possibility of Astra reaching a critical cyber capability threshold could not be ruled out, this monitoring obligation was expanded to include all inference work involving the use of Astra's tool. A key point of this change is that previously, the focus was on the deployment of frontier models and RL training runs, but it did not apply to all inference and training runs.

A Policy of Not Passing On Monitoring Costs to Customers

According to The Register, an OpenAI spokesperson explained that this 20% monitoring cost is for internal research and will not be passed on to customers. This is a statement of policy that the development team will absorb the costs of ensuring security internally, and it can be seen as an example of how investment in AI security does not necessarily translate directly into increased burdens on end users.

A Long-Term Move: "Rewriting" the Preparedness Framework

Among these announcements, what is particularly noteworthy for researchers is OpenAI's plan to fundamentally rewrite its risk management framework, the "Preparedness Framework." Much of this framework was formulated in 2023—a time when these risks were still purely theoretical.

An OpenAI researcher told reporters, "As we train more powerful models, we want to have a very high level of confidence that we understand the scope of their capabilities, have the ability to measure them, and meet higher standards of alignment (aligning AI's goals with human intentions)." This reflects the recognition that the risk management framework itself needs to be continuously updated to match the pace of model capability improvement.

The "Pacing" Debate Involving the Entire Industry

OpenAI's recent actions are not merely a decision made by an individual company, but are linked to industry-wide movements. The "Pacing the Frontier" letter, signed by several companies including OpenAI and Anthropic, called on the U.S. government to develop technical and institutional mechanisms to appropriately slow down AI before its capabilities surpass human understanding and control.

Interestingly, there are differences in the level of response even within the same industry. Anthropic, for example, has reportedly taken the position that a large-scale slowdown like OpenAI's is unnecessary. This difference in response likely reflects differences in each company's internal risk assessment and perception of the model's capabilities.

Points to Note from a Researcher's Perspective

The monitoring architecture recently released by OpenAI demonstrates a shift in the nature of AI model safety evaluation, moving from mere "post-hoc testing" to a "continuously operating monitoring infrastructure." The 20% increase in computational costs is significant, suggesting that this type of safety measure will become a crucial element in future AI development resource allocation.

The question of whether small research institutions and startups can develop this type of high-cost monitoring infrastructure at the same level remains a challenge for the industry as a whole. While frontier companies' safety measures are becoming more sophisticated, we will continue to closely monitor how the benefits and cost burdens of these advancements will spread throughout the industry.

OpenAIAI安全性監視システムAIガバナンスAI/ML

A 600% surge on its first day of listing—the dancing robot company's stark contrast between "frenzy" and "reality"

On August 19th, Unitree Robotics recorded a 630% surge in its first day of trading on the Shanghai STAR market (up 460% at closing), reaching a market capitalization of approximately $50 billion and exhibiting record-breaking overheating with an application rate of 8,000 times. This article skeptically examines the gap between its size (only 480 employees) and its impressive performance, the fact that 42% of its 2025 sales will come from quadruped robots, the reality that industrial applications account for less than 10% of sales with demand primarily from research and educational institutions, and the evaluation by Counterpoint Research and expert criticisms of its technological limitations.

600% Surge on First Day of Listing—The Gap Between "Frenzy" and "Reality" Revealed by the Dancing Robot Company

On August 19th, Unitree Robotics' stock price surged to nearly 630% immediately after trading began on the Shanghai STAR Market, ultimately closing 460% higher. Its market capitalization is estimated to have reached approximately $50 billion. While we previously discussed the determination of the listing price, the actual listing results far exceeded expectations. As a software engineer, I want to examine the nature of this frenzy and the gap between it and the reality behind it.

The Overheating Revealed by an "8,000-Times" Application Ratio

The first thing that surprises with this IPO is the application ratio. Unitree's IPO saw an application ratio exceeding 8,000 times, a record level in the history of the Shanghai tech market, "STAR Market." Despite a 3% decline in China's overall benchmark index that day, Unitree's stock surged against the trend, demonstrating how investor expectations for this stock were unique and detached from the overall market sentiment.

The initial price increase was also outstanding. While the average first-day increase for newly listed Chinese companies this year was 279%, Unitree significantly surpassed this, recording a 460% increase (reaching nearly 630% at one point). The backing from prominent companies, such as DeepSeek (approximately 140 million yuan) and Tencent (an existing shareholder), is also seen as a contributing factor to this frenzy.

The Gap in Scale: "480 People" and "World's Largest"

What personally struck me most in this news was the scale of Unitree. Despite being valued at $50 billion in market capitalization, it reportedly has only around 480 employees. This is an extremely unusual ratio compared to the conventional wisdom of the manufacturing industry.

However, this small team structure itself doesn't necessarily need to be viewed negatively, as it reflects a software-centric business model. But when looking at these figures, it's necessary to carefully consider whether investors truly understand the reality of the "world's largest humanoid robot manufacturer" before buying shares, or whether they're simply reacting to the enthusiastic image associated with the terms "humanoid" and "AI."

42% of Sales are Actually from Robot Dogs

There's another fact often overlooked behind the headlines. Unitree's 2025 sales are projected to be approximately 170 million yuan (about $252 million), representing more than a tenfold increase in two years. However, a breakdown of these sales reveals that 42% comes not from bipedal humanoid robots, but from quadrupedal "robot dogs."

This means that the actual revenue structure of this company, often portrayed with the glamorous image of humanoid robots, still heavily relies on the more mature product category of quadrupedal robots. While this isn't necessarily a negative fact, it highlights the dangers of evaluating Unitree's stock based solely on the simplistic equation of "Unitree = humanoid company."

The Reality: Industrial Applications are "Still Less Than 10%"

Even more important is the breakdown of sales by actual use, as detailed in Unitree's prospectus. It reports that industrial applications accounted for less than 10% of sales in the first three quarters of 2025. The majority of sales come from research and educational institutions.

This directly relates to the theme of "between announcement and verified" discussed in a previous article. While the title of "world's largest humanoid manufacturer" is true based on shipment figures (over 5,500 units), the reality that the majority of these are not "robots actually working in factories," but rather "products used in laboratories or purchased as collector's items," is a point that investors often overlook.

Analyst Assessments and Remaining Technical Hurdles

Analysts at market analysis firm Counterpoint Research have assessed Unitree's IPO as "a significant milestone for the humanoid industry and a potential benchmark for future IPOs." Indeed, Chinese companies accounted for the majority of global humanoid robot shipments last year, significantly outpacing competitors such as Tesla, Figure AI, and Boston Dynamics in terms of volume.

On the other hand, experts point out that this technology still faces significant limitations in both hardware and software. The industry consensus is that it will take several more years before bipedal robots are fully deployed on factory lines and in homes.

What Software Professionals Should Consider

This IPO does not negate Unitree's technological capabilities or market position. Its track record of over 5,500 humanoid shipments and over 33,000 quadruped robot cumulative sales demonstrates a solid position in this field. However, there is a significant gap between the "$50 billion valuation" that is at the heart of investor enthusiasm and the reality of being a "company of 480 people," with "42% of sales coming from robotic dogs" and "less than 10% in industrial applications."

Going forward, it will be necessary to continuously examine the company's capabilities through more detailed financial information disclosed as a publicly listed company, to see how the funds raised through this IPO will actually be invested in research and development and expanding production capacity, and to what extent the proportion of industrial applications can be increased.

UnitreeヒューマノイドIPO中国株式市場
Advertisement300 × 250