The Day the "Intelligence" vs. "Speed" Dichotomy Ends—OpenAI and Cerebras Achieve 750 Tokens/Second
On August 13th, OpenAI released a preview of its new API service tier, "Ultrafast." This allows their top-of-the-line model, "GPT-5.6 Sol," to run up to 14 times faster than standard processing, achieving a maximum speed of 750 tokens per second. What's particularly interesting from an engineer's perspective is that this speed increase isn't achieved by making the model smaller and lighter, but through collaboration with a dedicated hardware partner.
An Approach That Overturns Conventional Wisdom
Until now, developers in the AI industry seeking near real-time response speeds effectively had only one option: a compromise to gain response speed by choosing smaller, less powerful models. According to OpenAI itself, Ultrafast is the first attempt to break this long-standing "speed or intelligence" trade-off.
Ultrafast is a mechanism that runs the existing top-of-the-line GPT-5.6 Sol model on faster hardware; it's not a new model with modified internal workings. In other words, the idea is to "maintain the same intelligence while only changing the speed provided."
Cerebras's "Wafer-Scale Chip" as a Trump Card
This speed increase is supported by a special semiconductor called a "wafer-scale chip," developed by a company called Cerebras. Normally, semiconductor chips are manufactured by cutting numerous small chips from a disc-shaped silicon wafer. However, Cerebras' approach is completely different: they use an entire wafer as one giant chip without cutting it.
This design significantly reduces the latency associated with "data transfer between chips," which is unavoidable in typical GPU clusters. AI model inference processing requires frequent exchange of large amounts of parameter data between the computing unit and memory, and the speed of this data transfer often becomes a bottleneck in the final response speed. Cerebras' chip takes an approach that eliminates this physical constraint itself at the design stage.
Understanding the "14x" Speed with Concrete Numbers
According to comparative data released by OpenAI and Cerebras, Ultrafast is 5 times faster than Anthropic's "Claude Opus 4.8" in Fast mode and 11 times faster than "Claude Fable 5" (based on output speed data for each model reported by Artificial Analysis).
As a further example demonstrating its practical effectiveness, it has been reported that GPT-5.6 Sol Ultrafast processed all questions in the "Humanity's Last Exam," a graduate-level benchmark spanning chemistry, economics, and literature consisting of 2,500 questions, in just over 11 hours. Standard processing would have taken significantly longer for the same task.
Expected Use Cases: Tasks Where You Can't Wait
OpenAI cites incident response (emergency response to system failures) as an example of its application. By analyzing logs, code change histories, and reports in real time while a failure is occurring, it's possible to dramatically increase the speed at which engineers can identify the cause and prepare fixes.
Besides this, applications are envisioned in areas where response speed directly impacts user experience, such as coding assistance, commercial transactions, financial research, and customer support. Jane Street's AI assistant development team commented, "The speed improvements from Cerebras are impressive, enabling a new way of working where developers can collaborate with models while maintaining greater focus."
Another Significance: "Dual-Layering from NVIDIA Dependence"
The offering of Ultrafast has another strategic significance beyond simply competing on speed. OpenAI has already announced the establishment of its own AI chip development team, but this expanded partnership with Cerebras is positioned as part of a "multi-vendor strategy" to reduce its dependence on NVIDIA and have multiple hardware partners.
Cerebras and OpenAI have reportedly had a $10 billion partnership, and this Ultrafast announcement can be seen as a further step in that relationship.
Reservations Behind the "Up to" Expression
However, when evaluating this announcement, several reservations should be kept in mind. The figures of "14x" and "750 tokens/second" presented by OpenAI are both accompanied by the expression "up to," meaning that not all requests will always maintain this speed. Pricing and the official date of general availability have not yet been announced.
Currently, it is in a preview phase for a limited number of customers, and developers who actually use this service will need to measure the actual latency, output speed, reliability, and cost on their workloads before deciding to adopt it.
What Engineers Should Consider
Ultrafast demonstrates a shift in the competitive landscape for AI models, moving from simple "intelligence" to infrastructure-level competition focused on "how quickly and practically that intelligence can be delivered."
For developers who have previously had to choose smaller models for applications requiring real-time performance, this type of high-speed inference service has the potential to greatly expand their design options. We will be closely watching for further updates on how the pricing structure and general availability will be determined, and to what extent stability will be guaranteed in actual production environments.