Wednesday, August 26, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Evening Edition Free of Charge
TODAY'S LEAD STORY

"34 out of 36 articles were plagiarized"—OpenAI exposes the contents of a fake think tank created with AI.

On August 25, OpenAI announced it had suspended a group of ChatGPT accounts linked to clandestine operations originating from Russia. This article explains the reality that 34 out of 36 articles published by the non-existent expert think tank "International Burke Institute" were plagiarized, the concealment tactics including VPN access and instructions to erase linguistic traces from Russian language prompts, the unique "sovereignty index" that portrays Russia favorably, the three-category classification on the Brookings Breakout Scale, and the continuity with past cases such as the Rybar crackdown in February.

"34 out of 36 articles were plagiarized"—OpenAI exposes the contents of a fake think tank created with AI.
(Photo: illustrative)

"34 out of 36 posts were plagiarized"—OpenAI exposes the contents of a fake think tank created with AI

On August 25th, OpenAI announced that it had uncovered a covert operation originating from Russia and suspended a group of related ChatGPT accounts. The operation involved creating an organization posing as a non-existent think tank, using AI to generate a large amount of content, and spreading pro-Russian rhetoric. What's interesting from an engineer's perspective is the specific methods used in this operation and the investigation process that OpenAI used to uncover it.

An unexpected discovery that began with an "investigation of AI-generated content"

The trigger for this uncovering was OpenAI's own investigation of AI-generated social media posts. In the process, the existence of a far larger and more organized operation than initially anticipated was revealed.

At the heart of the operation was the "International Burke Institute (IBI)," a self-proclaimed "expert community" based in Israel. Registered in February 2025, this website claimed to have a wide range of experts, including real, prominent scholars such as Francis Fukuyama and Noam Chomsky.

Plagiarism: 34 out of 36 articles

OpenAI extracted and examined 36 expert-related articles published on the IBI website and found that 34 of them were direct copies from other parts of the internet. In some cases, the articles were even attributed to the wrong authors.

In other words, this "think tank" was essentially disguising non-existent expertise as if it were real through plagiarized content. Furthermore, the organization used its own evaluation metric, the "sovereignty index," which was an arbitrary scoring system designed to portray Russia favorably and criticize Western countries. Reports on France, Germany, the United States, and the EU, according to OpenAI, "touched controversial territory."

A Double Layer of Concealment: VPNs and "Erasing Traces"

From a technical standpoint, the methods used by those carrying out this operation to conceal their identities are particularly interesting. Because OpenAI does not allow access to its models from Russia, the operators accessed the service via a VPN (Virtual Private Network, a technology that disguises the actual connection source).

Furthermore, they followed a procedure of inputting prompts in Russian and then generating English social media content based on that input. They explicitly instructed ChatGPT to intentionally remove linguistic features that would suggest the user was a Russian speaker. The generated content was posted to multiple platforms, including Substack, Telegram, X, Facebook, and LinkedIn.

An Assessment Where "Infrastructure" is Valued More Than "Appeal"

What is particularly insightful in OpenAI's analysis is its assessment of the actual impact of this operation. The company states that, in terms of the number of readers actually reached, the impact was limited.

However, OpenAI positions the importance of this operation not in the "scale of the readership reached," but in the "infrastructure itself that was built." The method of creating a seemingly trustworthy organization with fake experts, reprinted academic research, and a unique (or disguised) risk index could be reused as a foundation for larger-scale covert operations in the future. OpenAI has classified this operation as Category 3 on its proprietary "Brookings Breakout Scale" classification system.

Similar Busts in the Past

This is not the first such incident for OpenAI. In February, the company announced the suspension of a ChatGPT account linked to Rybar, a pro-Russian media company that the UK government described as "partially colluding with the Russian Presidential Administration." In June 2025, it also busted four Chinese-based covert operations operating across platforms such as TikTok, Facebook, Reddit, and X.

The fact that this type of crackdown is occurring repeatedly indicates that information manipulation using AI-generated content is not a one-off incident, but a persistent threat.

What Engineers Should Consider

This incident demonstrates the reality that AI-generated content dramatically eliminates the "content creation bottleneck" in information manipulation. Generating large amounts of plausible content, which previously required the mobilization of numerous human writers and copywriters, is now possible with just a small number of operators and access to AI models.

On the other hand, the fact that this case was discovered and exposed by OpenAI itself indicates that the monitoring and detection capabilities of the platform are also continuously evolving. For engineers operating platforms that handle AI-generated content, creating mechanisms to detect this type of "organized abuse pattern" will likely become an increasingly important design requirement.

OpenAIChatGPT情報工作AI安全性サイバーセキュリティ

"21% of the text was AI-generated"—The reality of the collapse of the peer-review system faced by the largest AI research conference.

AI detection company Pangram Labs analyzed all 76,139 peer-reviewed reports from ICLR 2026 and determined that 21% (approximately 15,899 reports) were fully AI-generated. This study outlines the research methodology, including the reasons behind the 2.7-fold increase in the number of submissions in three years from 2024, the correlation between AI-generated reviews giving higher scores to low-quality papers, the "Hivemind effect" showing statistical significance of p<0.05 in all 21 fields, independent verification that 50 out of 58 author complaints (86.2%) were consistent with Pangram's judgment, and the discrepancy in detection rates with Sem-Detect (21% vs. 5%).

"21% Fully AI-Generated"—The Collapse of the Peer-Review System Faced by the Largest AI Research Conference

Pangram Labs, an AI detection company, has published the results of its analysis of all 76,139 peer-review reports submitted to ICLR 2026, one of the largest international conferences in the field of AI. Of these, 21%, or approximately 15,899, were determined to be entirely AI-generated. As a journalist with a background in AI research, I want to carefully examine the reality and background of this ironic situation where "AI is reviewing AI research."

From "7,304" to "19,814"—Submissions Nearly Triple in Three Years

Behind this problem lies the explosive increase in the number of submissions to ICLR itself. The number of submitted papers, which was 7,304 in 2024, ballooned to 11,672 in 2025 and 19,814 in 2026—a roughly 2.7-fold increase in just three years.

To handle this surge in submissions, ICLR adopted an operational system of securing additional reviewers from the author's own pool. As a result, situations arose where undergraduate students were handling reviews alongside professors, and each reviewer was assigned multiple papers under tight deadlines. Ironically, this "system designed to handle volume" has led to a "mass-production approach" in the form of AI-generated reviews.

Several "Characteristics" Common to AI-Generated Reviews

According to Pangram Labs' analysis, review reports identified as AI-generated shared several common characteristics. Typical examples include long sentences, frequent use of bold headings, and low information density. One peer review report, reportedly 3,000 words long, listed 40 weaknesses and 40 questions. Researchers described this as "extensive but lacking in substance."

Even more noteworthy is the finding that AI-generated reviews correlated with higher scores, regardless of the actual quality of the papers. This suggests that reviewers weren't simply using AI as an aid, but were entrusting the entire decision-making process to the AI.

Statistically Verified "Hivemind Effect"

Another study, "Stop Automating Peer Review Without Rigorous Evaluation," which further analyzed this issue, reported an interesting statistical finding: AI-generated reviews showed a statistically significantly higher tendency to be similar in content to human-written reviews—a "Hivemind effect" (herd effect).

This effect was shown to be statistically significant (p<0.05) across all 21 major areas of the ICLR. The effect size (a statistical indicator showing the magnitude of the difference between different groups) becomes even larger when focusing specifically on the "weaknesses" and "questions" sections. In other words, the data supports the tendency for AI-generated peer review comments, unlike the diverse perspectives produced by humans, to converge on similar, uniform criticisms.

Independent Verification through "Author Complaints"

To evaluate the reliability of this analysis, verification from another angle is helpful. The research team searched through all 159,775 author comments from ICLR 2026, identifying 58 cases where authors claimed the reviews were AI-generated.

Of these, Pangram Labs determined 50 cases (86.2%) to be "fully AI-generated," while only 2 cases (3.4%) were determined to be "fully human-written." This provides independent evidence indicating a high degree of agreement between the authors' subjective doubts and the objective judgments of the AI ​​detection tool.

ICLR's Response and Remaining Issues

Following this issue, two weeks prior to Pangram Labs' announcement, ICLR had already announced rejection measures for papers containing undisclosed LLM usage and disciplinary action against reviewers who submitted "reviews containing hallucinations." Regarding references containing hallucinations (a phenomenon where AI generates plausible information that is not based on facts), another detection tool, GPTZero, reportedly detected more than 50 such cases out of a mere 300 samples.

However, caution is needed when interpreting these detection results. A comparative analysis using another detection method, "Sem-Detect," showed a different distribution (21% AI-generated compared to 5% with Sem-Detect) for the same ICLR 2026 data, indicating that the accuracy and assumptions of the detection method itself remain debatable.

What Researchers Should Consider

This series of events highlights a structural irony: the AI ​​research community itself is having the reliability of its core peer-review system threatened by the very technology it created. While the number of submitted papers is increasing exponentially, the number of human reviewers cannot keep pace. AI is unintentionally infiltrating this supply-demand gap by filling it.

In the future, alongside improvements in the accuracy of AI detection technology, the peer-review process itself will need to be redesigned to accommodate the presence of AI. Where and how should we draw the line between human and AI roles? This question is not limited to the ICLR conference; it's an unavoidable challenge for the entire academic publishing ecosystem.

査読AI検出学術研究ICLRAI/ML

"He surpassed Usain Bolt," then he fell and caught fire—the true meaning behind Beijing's record-breaking performance.

On August 25th, at the 2nd World Humanoid Robot Competition in Beijing, a Chinese-made robot set a new human record in the 100m dash with a time of 8.86 seconds, but it toppled over and caught fire immediately after crossing the finish line. This article skeptically examines the rapid record improvement from 9.39 seconds just three days earlier, the candid statement by the People's Daily itself that "housekeeper and caregiver robots are still in the pilot stage," and the fundamental gap between competitions and mass production pointed out by Digitimes.

"Surpassing Usain Bolt," Then Falling and Sparking—The True Meaning Behind Beijing's Record-Breaking Event

On August 25th, at the 2nd World Humanoid Robot Competition held in Beijing, a Chinese-made humanoid robot set a new record in the 100-meter dash, clocking in at 8.86 seconds. This record surpasses the 9.58-second world record for men's 100 meters, set by Usain Bolt in 2009. However, the robot that set this record fell immediately after crossing the finish line, crashing into a padded barrier with sparks flying from its waist. As a software engineer, I want to examine the current state of the industry as symbolized by this scene.

Rapid Record-Breaking: "0.53 Seconds in 3 Days"

Before this record, a record of 9.39 seconds had already been set just three days earlier, on August 22nd. This means that records were being rapidly broken even within the short period of the competition.

This "speed" itself is certainly a remarkable advance. However, it's important to calmly consider that a single, linear metric like a 100-meter dash cannot adequately measure the overall practicality of humanoid robots. In robotics, "running fast" and "moving safely and stably while maintaining balance in complex and unpredictable environments" are entirely different technical challenges. The outcome of this incident—a fall and subsequent fire—ironically demonstrates the significant gap that still exists between these two challenges.

China's National Strategy: An "Olympic-Style Competition"

This World Humanoid Robot Competition is now in its second year, having held its first event in 2025. It is held at the National Speed ​​Skating Hall in Beijing (a facility built for the 2022 Beijing Winter Olympics) and is conducted in a format similar to the Olympics.

This event is not merely a technology demonstration; it is positioned as part of a clear national strategy to showcase China's rapid progress in robotics to both domestic and international audiences amidst the intensifying technological competition with the United States. It can be seen as a culmination of the massive financial and policy support that the Chinese government has repeatedly invested in this field.

Candid Assessment by Chinese State Media

What I personally find most noteworthy in this series of reports is the candid assessment offered by the People's Daily, the official newspaper of the Communist Party of China. Prior to the competition, the newspaper stated that "robots as 'housekeepers' and 'caregivers' are still in the pilot and verification stages," adding that "(however) these challenges are precisely the focus of development... China's robotics industry is evolving towards higher quality and innovation."

This is a valuable statement from China itself, acknowledging that, behind the "showpieces" of record-breaking performances at spectacular competitions, numerous fundamental challenges remain in the pursuit of practical application. The fact that the host country of an event of national prestige would choose to publicly express such caution is itself important information for evaluating the reality of this field.

The Next Hurdle: Mass Production

Digitimes' analysis, based on on-site coverage, accurately articulates the more fundamental question raised by this competition. Behind the glamorous aspects of record-breaking sprints and human-robot sports demonstrations, the real question facing this industry is whether "humanoid robots can move beyond the showcase stage and reach the practical application stage on actual production lines."

Considering the current situation where companies like Unitree and XPeng, which we have covered previously, are struggling to establish mass production systems, this observation is spot-on. A significant technical and manufacturing gap still exists between the spectacular performances seen in competitions and the practical, reliable operation they offer in factories and homes.

What Software Engineers Should Consider

The recent record-breaking 100-meter dash, followed immediately by the fall and fire, symbolically illustrates the current state of this industry. While robots can already surpass human capabilities in specific, narrow tasks (in this case, running in a straight line), they still lack reliability in essential, practical actions such as stable stopping and deceleration.

When evaluating the capabilities of humanoid robots, it's crucial to consider not only the impressive record numbers but also the conditions under which those records were achieved and their reproducibility. As the People's Daily itself acknowledges, the real test for this industry is yet to come, outside the arena—in actual homes, care facilities, and factories.

ヒューマノイド中国北京競技会フィジカルAI
Advertisement300 × 250