Tuesday, August 18, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Evening Edition Free of Charge
TODAY'S LEAD STORY

The day the dichotomy between "intelligence" and "speed" ends—the details of the 750 tokens/second achieved by OpenAI and Cerebras.

On August 13th, OpenAI released a preview of its new API tier, "Ultrafast," which runs GPT-5.6 Sol at up to 14 times faster and 750 tokens per second. The article explains how it reduces data transfer latency using Cerebras wafer-scale chips, compares speeds of 5 times faster than Claude Opus 4.8 and 11 times faster than Claude Fable 5, provides a real-world example of processing Humanity's Last Exam in just over 11 hours, discusses anticipated applications such as incident response, and explains its significance as a multi-vendor strategy to reduce reliance on NVIDIA.

The day the dichotomy between "intelligence" and "speed" ends—the details of the 750 tokens/second achieved by OpenAI and Cerebras.
(Photo: illustrative)

The Day the "Intelligence" vs. "Speed" Dichotomy Ends—OpenAI and Cerebras Achieve 750 Tokens/Second

On August 13th, OpenAI released a preview of its new API service tier, "Ultrafast." This allows their top-of-the-line model, "GPT-5.6 Sol," to run up to 14 times faster than standard processing, achieving a maximum speed of 750 tokens per second. What's particularly interesting from an engineer's perspective is that this speed increase isn't achieved by making the model smaller and lighter, but through collaboration with a dedicated hardware partner.

An Approach That Overturns Conventional Wisdom

Until now, developers in the AI ​​industry seeking near real-time response speeds effectively had only one option: a compromise to gain response speed by choosing smaller, less powerful models. According to OpenAI itself, Ultrafast is the first attempt to break this long-standing "speed or intelligence" trade-off.

Ultrafast is a mechanism that runs the existing top-of-the-line GPT-5.6 Sol model on faster hardware; it's not a new model with modified internal workings. In other words, the idea is to "maintain the same intelligence while only changing the speed provided."

Cerebras's "Wafer-Scale Chip" as a Trump Card

This speed increase is supported by a special semiconductor called a "wafer-scale chip," developed by a company called Cerebras. Normally, semiconductor chips are manufactured by cutting numerous small chips from a disc-shaped silicon wafer. However, Cerebras' approach is completely different: they use an entire wafer as one giant chip without cutting it.

This design significantly reduces the latency associated with "data transfer between chips," which is unavoidable in typical GPU clusters. AI model inference processing requires frequent exchange of large amounts of parameter data between the computing unit and memory, and the speed of this data transfer often becomes a bottleneck in the final response speed. Cerebras' chip takes an approach that eliminates this physical constraint itself at the design stage.

Understanding the "14x" Speed ​​with Concrete Numbers

According to comparative data released by OpenAI and Cerebras, Ultrafast is 5 times faster than Anthropic's "Claude Opus 4.8" in Fast mode and 11 times faster than "Claude Fable 5" (based on output speed data for each model reported by Artificial Analysis).

As a further example demonstrating its practical effectiveness, it has been reported that GPT-5.6 Sol Ultrafast processed all questions in the "Humanity's Last Exam," a graduate-level benchmark spanning chemistry, economics, and literature consisting of 2,500 questions, in just over 11 hours. Standard processing would have taken significantly longer for the same task.

Expected Use Cases: Tasks Where You Can't Wait

OpenAI cites incident response (emergency response to system failures) as an example of its application. By analyzing logs, code change histories, and reports in real time while a failure is occurring, it's possible to dramatically increase the speed at which engineers can identify the cause and prepare fixes.

Besides this, applications are envisioned in areas where response speed directly impacts user experience, such as coding assistance, commercial transactions, financial research, and customer support. Jane Street's AI assistant development team commented, "The speed improvements from Cerebras are impressive, enabling a new way of working where developers can collaborate with models while maintaining greater focus."

Another Significance: "Dual-Layering from NVIDIA Dependence"

The offering of Ultrafast has another strategic significance beyond simply competing on speed. OpenAI has already announced the establishment of its own AI chip development team, but this expanded partnership with Cerebras is positioned as part of a "multi-vendor strategy" to reduce its dependence on NVIDIA and have multiple hardware partners.

Cerebras and OpenAI have reportedly had a $10 billion partnership, and this Ultrafast announcement can be seen as a further step in that relationship.

Reservations Behind the "Up to" Expression

However, when evaluating this announcement, several reservations should be kept in mind. The figures of "14x" and "750 tokens/second" presented by OpenAI are both accompanied by the expression "up to," meaning that not all requests will always maintain this speed. Pricing and the official date of general availability have not yet been announced.

Currently, it is in a preview phase for a limited number of customers, and developers who actually use this service will need to measure the actual latency, output speed, reliability, and cost on their workloads before deciding to adopt it.

What Engineers Should Consider

Ultrafast demonstrates a shift in the competitive landscape for AI models, moving from simple "intelligence" to infrastructure-level competition focused on "how quickly and practically that intelligence can be delivered."

For developers who have previously had to choose smaller models for applications requiring real-time performance, this type of high-speed inference service has the potential to greatly expand their design options. We will be closely watching for further updates on how the pricing structure and general availability will be determined, and to what extent stability will be guaranteed in actual production environments.

OpenAICerebras推論高速化半導体AI/ML

Why is China's autonomous driving technology spreading from Zagreb to Europe? – An analysis of Pony.ai and Uber's 2,000-vehicle plan.

On August 14th, Pony.ai and Uber announced an expansion of their partnership, adding four more European cities to their existing service in Zagreb, bringing the total to over 2,000 robotaxis, with plans for expansion into the Middle East. This article skeptically examines the three-tiered division of labor model—technology, platform, and local operations—the backdrop of China's suspension of robotaxi license issuance, the rush of European expansion by competitors such as WeRide and Momenta, and the dangers of directly extrapolating Chinese domestic performance to European regulations and traffic conditions.

Why is Chinese-developed autonomous driving technology spreading from Zagreb to Europe? – An analysis of Pony.ai and Uber's 2,000-robotaxi plan

On August 14th, Chinese autonomous driving company Pony.ai and ride-hailing giant Uber announced an expansion of their partnership in Europe. Starting with their existing commercial robotaxi service in Zagreb, Croatia, they plan to add four more European cities, deploying a total of over 2,000 robotaxi vehicles. Plans for expansion into the Middle East were also announced simultaneously. As a software engineer, I want to examine the technical and structural implications of this partnership.

A clever division of labor model: "We won't deploy it ourselves"

What's interesting about the Pony.ai and Uber partnership is the clear division of labor between the two companies. Pony.ai provides SAE Level 4 technology (a level of autonomous driving where fully autonomous driving is possible under specific conditions, and human intervention in emergencies is generally unnecessary) and insights gained from existing large-scale fleet operations. Uber provides a mobility platform already established worldwide, encompassing booking, payment, and customer support. Daily vehicle operations are handled by local partners selected for each market.

In the Zagreb example, the Croatian mobility company Verne is responsible for vehicle ownership and daily operations. This three-tiered structure of "technology provider, platform provider, and local operator" eliminates the need for companies with autonomous driving technology to establish local subsidiaries from scratch in each country and city, significantly accelerating deployment.

The "Europe" Option for Chinese Companies

Pony.ai already has a track record of deploying fully autonomous robotaxi services without human drivers in major Chinese cities such as Beijing, Shanghai, Guangzhou, and Shenzhen. Bringing this experience to Europe can be seen as a strategy by Chinese autonomous driving companies to supplement growth opportunities that cannot be met by the domestic market alone through overseas expansion.

What's interesting is that this partnership was announced against the backdrop of a several-month suspension of robotaxi license issuance in China due to industry-wide safety reviews. It's reported that these reviews were completed at the end of June, and license issuance has begun to resume in some cities. Given the temporary uncertainty in the domestic regulatory environment, actively expanding into overseas markets makes sense as a risk diversification strategy for companies.

Points to Consider and Examine from a Skeptical Perspective

When evaluating news about this type of Chinese-developed autonomous driving technology, there are several points that require careful consideration. First, the announcement hasn't yet revealed the specific names of the "four cities" or the target date for achieving the "2,000 units" figure. Details are expected to be announced in stages, and at this point, it's primarily a statement of intent.

Furthermore, whether the operational performance in China can be replicated under European road conditions, traffic regulations, and increasingly stringent European AI and autonomous driving regulations requires a completely separate examination. As we've seen, Europe has a strict regulatory environment, including the AI ​​Act, and it would be premature to simply extrapolate performance and safety based solely on China's track record.

Accelerating Overseas Expansion of China-Originated Autonomous Driving Systems Across the Industry

Pony.ai's move is not an isolated case. Its competitor, WeRide, has partnered with Denmark's GreenMobility and plans to deploy autonomous vehicles in Denmark in 2027. Similarly, Momenta, which has partnered with Uber, has reportedly become the first Asian company to obtain permission to test robotaxis throughout Germany.

This situation, where Chinese autonomous driving companies are accelerating their expansion into the European market as if by prior arrangement, is likely a result of two factors: the increasing technological maturity within China and the uncertainty of the regulatory environment in the domestic market.

Things to Watch from a Software Perspective

The expanded partnership between Pony.ai and Uber demonstrates an effective business model for accelerating the international deployment of autonomous driving technology through a division of labor model involving technology provision, platform provision, and local operations. This mechanism itself is a sufficiently rational approach, both technically and strategically.

On the other hand, it's necessary to closely monitor future updates to see at what pace and in which cities the ambitious figure of "2,000 vehicles" will actually be realized. To what extent will the success in China be replicated in the different regulatory and traffic environment of Europe? Pony.ai is scheduled to announce its second-quarter results on August 18th, and the financial indicators and operational performance details presented there will likely be important factors in judging the feasibility of this plan.

自動運転ロボタクシーUber中国欧州

They were neck and neck in "finding," but far behind in "creating"—the complexity of capability assessment revealed by China's open model.

On August 14th, China's Z.ai announced that its GLM-5.3 had achieved a score of 84.5% on the CyberGym benchmark, slightly surpassing Anthropic's Mythos 5 (83.8%). This article explains the statistical limits of a difference of less than one percentage point, the crucial difference in vulnerability discovery and weaponization capabilities (54.4% vs. 78.0% on ExploitBench), the release of weights which will be withheld until August 28th, the field testing which discovered 2,436 vulnerabilities from 269 projects, and the technical background including post-training with an additional 743 billion parameters.

Matching in "Finding," but a Large Difference in "Creating"—The Complexity of Capability Evaluation Revealed by a Chinese Open Model

On August 14th, Chinese AI company Z.ai (formerly Zhipu AI) released its new model, "GLM-5.3." The company announced that this model scored slightly higher than Anthropic's restricted model, "Mythos 5," in the cybersecurity benchmark "CyberGym." As a journalist with a background in AI research, I would like to carefully examine what this "slightly higher" result actually means and what it doesn't.

What the CyberGym Benchmark Measures

CyberGym is a benchmark consisting of 1,507 tasks that evaluate whether an AI model can inspect source code, find security flaws, and confirm that those flaws are genuine vulnerabilities. According to Z.ai's announcement, GLM-5.3 achieved a score of 84.5% in this benchmark, slightly surpassing Mythos 5's 83.8% and OpenAI's GPT-5.6 Sol's 83.6%.

Z.ai has also released details of the evaluation conditions. Using the Claude Code 2.1.207 harness (an evaluation execution environment), with maximum inference effort set, no web tools were used, and evaluations were conducted with one trial (Pass@1) per task under the conditions of temperature parameter 1.0, top P 1.0, and maximum output length of 128,000 tokens.

Correctly Interpreting a "Slight Difference"

First, it's important to consider how statistically significant this 84.5% vs. 83.8% difference is. This difference of less than one point is easily exacerbated by slight differences in benchmark design and evaluation conditions, and is not significant enough to definitively conclude that "GLM-5.3 clearly outperformed Mythos 5."

Furthermore, it's important to note that these figures were all published by Z.ai itself and have not undergone independent verification by a third party. When comparing scores from competitors' models, it's crucial to keep in mind that differences in evaluation conditions can influence the results.

The Decisive Difference Between "Finding" and "Building" Vulnerabilities

What's most noteworthy for researchers in this announcement isn't the closeness of the CyberGym scores, but the far greater difference in the subsequent "ExploitBench" benchmark. ExploitBench is a more advanced test that evaluates whether discovered vulnerabilities can be transformed into working attack code (exploits).

In this benchmark, GLM-5.3 scored only 54.4%, while Mythos 5 achieved 78.0% and GPT-5.6 Sol 76.5%, all significantly outperforming GLM-5.3. Even in a more practical evaluation metric—a time-limited attack development task (how many out of 105 tasks could be completed within two hours)—Mythos 5 reportedly completed the equivalent of 181 tasks, further highlighting the difference with GLM-5.3.

In short, while there is little difference between the two in the relatively passive ability to "find vulnerabilities and determine their authenticity," a significant gap still exists in the more active and advanced ability to "actually weaponize those vulnerabilities." This clearly demonstrates the danger of judging the overall picture based on a single benchmark score when evaluating the cybersecurity capabilities of AI models.

The Significance of Open Weights

Another point worth noting is that GLM-5.3 is provided as an open weight model (a model with publicly available trained parameters that anyone can download and use). However, the actual release of the model's weights is scheduled to be withheld until around August 28th for security review.

While Mythos 5 is a model with strict access restrictions by Anthropic, essentially a "gatekeeper" model, if a model with comparable capabilities becomes openly distributed, the nature of the "dual-use" problem in cybersecurity research will change significantly. In fact, it has been reported that GLM-5.3, in field tests conducted by a security team in China, discovered 2,436 vulnerabilities from 269 open-source projects, of which 1,097 were classified as medium to high severity.

Technical Evolution from the Previous Version

GLM-5.3 is based on the same 743 billion parameter mixed expert architecture as its predecessor, GLM-5.2. However, instead of training a new model from scratch, its capabilities have been enhanced through additional post-training. According to Z.ai, coding performance has improved by approximately 50% compared to the previous version.

This reflects the recent trend in efficiency in AI development, where, once the underlying model architecture is established, subsequent performance improvements can be continuously achieved through relatively low-cost additional training.

Points for Researchers to Consider

Z.ai's announcement is easily headlined as "Chinese AI company rapidly catching up to the US Frontier Laboratory in the field of cyber defense." However, a careful breakdown reveals a more precise picture: while they are indeed close in relatively basic tasks like "vulnerability discovery and verification," a clear gap still exists in the more advanced task of "actually weaponizing those vulnerabilities."

When evaluating the capabilities of an AI model, it's important to not only look at a single benchmark score that might easily grab headlines, but also to examine what that benchmark actually measures and how it compares to other benchmarks. We hope that once the GLM-5.3 weights are officially released, more detailed verification results from independent researchers will emerge.

Z.aiサイバーセキュリティベンチマークオープンソースAIAI/ML
Advertisement300 × 250