AI "Accidentally Confesses" Its Weakness—An Unprecedented Discovery Method Exposes a Copilot Vulnerability
On August 18th, Microsoft released a patch to fix a serious vulnerability called "CoSnitch." This vulnerability allows attackers to silently steal data from a victim's linked accounts, such as email, cloud storage, and calendar, simply by clicking a single link. While the technical severity is noteworthy, what was most interesting to me as an engineer was how this vulnerability was discovered.
An Attack Chain Linking Three Vulnerabilities
CoSnitch (CVE-2026-24301) is an attack that combines three different vulnerabilities present in Microsoft Copilot Personal. The first is the combination of Copilot's standard query parameter "?q=" with an undocumented, undisclosed parameter, allowing the attacker to automatically execute a prompt upon page loading, without any clicks or confirmations.
The second method involves using these automatically executed prompts to retrieve data from OAuth applications (such as Gmail, Google Drive, and Calendar) that the user has already granted access to Copilot, and then silently sending it to a server prepared by the attacker. It's crucial to note that this is not "unauthorized OAuth access." The victim had already granted Copilot access to these linked applications. CoSnitch broke the implicit assumption that "operations on linked applications only occur at the user's intentional request."
The third method is an indirect prompt injection (an attack technique that injects malicious instructions into normal content) that exploits the web summarization feature. The attacker prepares a seemingly harmless webpage by embedding malicious instructions in HTML comments, metadata, or visually hidden text. When a victim requested a summary of this page from Copilot, Copilot processed not only the displayed text but also hidden attacker instructions, resulting in malicious instructions being written to the user's "persistent memory."
Discovered New Technique: "Meta-Hacking"
The technique used by security company Varonis Threat Labs, which discovered this vulnerability, is attracting attention in the industry. The company calls this "meta-hacking." While this type of vulnerability is usually discovered through reverse engineering of the code, Varonis researchers took an approach of repeatedly asking Copilot "why automated execution is impossible."
Each time Copilot refused to answer the question, the explanation for the refusal included information about the system's technical workings. The researchers restructured each refusal into further questions, gradually refining the accuracy of the information that could be extracted from Copilot. In the process, Copilot itself, unsolicited, disclosed undisclosed URL parameters that would be used in the attack, midway through its explanation of the refusal. Varonis described this series of interactions as "a combination of highly sophisticated social engineering against LLMs, various jailbreaks, and prompt injection attacks."
AI Assistants Create Their Own "Map of Attack Surfaces"
This meta-hacking technique demonstrates the risk that the very act of an AI assistant attempting to explain its own security can unintentionally provide attackers with useful information. Refusal messages like "This cannot be done" are often accompanied by technical reasons for their inability. As these explanations accumulate, attackers can reconstruct the system's internal structure through dialogue alone, without ever examining the code.
This represents a new type of risk, distinct from conventional software security practices. While traditional bug detection involved analyzing static code and communication packets, in the case of AI assistants, clues to vulnerabilities can be extracted by "conversing" with the system itself.
A Third Vulnerability in Copilot
For Varonis, CoSnitch marks the third Copilot-related vulnerability discovery this year. In the past, Varonis has reported "Reprompt," which bypasses security mechanisms by repeatedly asking the same question, and "SearchLeak," which turns Microsoft 365 Copilot Enterprise into a covert data exfiltration route. All three vulnerabilities share an extremely simple starting point for attacks: "single click on a seemingly normal link."
Varonis reported this issue to Microsoft in December 2025, and it took approximately eight months for the actual patch to be released. Fortunately, no evidence of this vulnerability being exploited before the patch was released has been found.
What Engineers Should Consider
The lesson this news teaches is that careful auditing is necessary when designing AI assistants to have access to a wide range of data, such as email, calendars, and cloud storage. Varonis recommends that corporate security teams regularly audit third-party applications connected to Copilot, treat AI assistants as "authorized insiders requiring the same level of access oversight as human employees," and establish monitoring systems to detect anomalous data access via AI assistants. When integrating an AI assistant into your company's development workflow, it's worth taking stock of how much data access you're granting, and under what conditions, before considering its "convenience."