Models

OpenAI Just Broke the AI Triangle. Speed and Intelligence No Longer Trade Off.

CRAZE CRAZE Summary 3 things to know
  • Ultrafast runs GPT-5.6 Sol up to 14× faster with no intelligence loss, ending the speed-vs-smarts trade-off.
  • Cerebras' on-chip SRAM eliminates the GPU memory bandwidth bottleneck, enabling real-time frontier AI for critical-path workflows.
  • OpenAI's $10B Cerebras partnership validates non-GPU inference, pressuring Nvidia and AMD to accelerate alternative architectures.
Jeff Editorial | · 5 min read
OpenAI Just Broke the AI Triangle. Speed and Intelligence No Longer Trade Off.

On August 13, OpenAI previewed Ultrafast, a new service tier for GPT-5.6 Sol that runs up to 14× faster than Standard processing—up to 750 output tokens per second, with no loss in intelligence. Until now, getting real-time speed meant choosing a smaller model. That trade-off is over.

The announcement is brief. Ultrafast is powered by Cerebras' Wafer-Scale Engine architecture, launching first in the OpenAI API in limited preview to a small group of customers. But the signal it sends is larger than the product itself.

OpenAI's framing is direct: "Until now, getting real-time speed typically meant choosing a smaller or more specialized model." Ultrafast points in a new direction: more useful work per second, at frontier intelligence levels.

Ultrafast Is 14× Faster—and Just as Smart

The benchmark numbers are stark. On Humanity's Last Exam—2,500 graduate-level questions spanning chemistry, economics, and literature—GPT-5.6 Sol Ultrafast completed the full set in just over 11 hours. Claude Fable 5 needed 78 hours. Same accuracy, nearly 7× faster.

On GDP-Val, a benchmark of economically valuable knowledge-work tasks—legal briefs, financial models, engineering reports—Ultrafast delivered a 5.6× end-to-end speedup with no quality loss.

The competitive gap is also wide. According to output speed data from Artificial Analysis, Ultrafast is 5× faster than Claude Opus 4.8 in Fast mode, and 11× faster than Claude Fable 5. For developers building agentic workflows where latency compounds across multiple turns, that gap translates directly into productivity.

OpenAI Just Broke the AI Triangle. Speed and Intelligence No Longer Trade Off.
GPT‑5.6 Sol Ultrafast and standard build a working 3D warehouse simulator from the same text prompt, side by side.

44GB of SRAM Killed the Memory Wall

The speed comes from Cerebras' unique hardware design. Traditional GPU inference is constrained by memory bandwidth: model weights are stored in HBM memory and shuttled to compute units for every token generation step. That "memory wall" is the bottleneck.

Cerebras eliminates it. The Wafer-Scale Engine keeps 44 GB of model weights entirely on-chip in SRAM—orders of magnitude faster than HBM, with a reported 21 petabytes per second of memory bandwidth. The GPU's bottleneck becomes Cerebras' feature.

This architecture is why Cerebras has become OpenAI's strategic partner. In January 2026, OpenAI announced a deal to deploy 750 megawatts of Cerebras accelerators through 2028, valued at over $10 billion. Ultrafast is the first product of that partnership.

AI Just Entered the Critical Path

The most important implication is not about speed. It is about where AI can be deployed.

OpenAI's internal use cases reveal the shift. Engineering teams are using Ultrafast during active system outages: reading logs, analyzing traces, aggregating conversations, and helping prepare remediation plans while the outage is still unfolding. Research teams that used to launch experiments overnight and check results in the morning are now iterating multiple times within a single working day.

This changes the economics of AI adoption. Until now, AI agents were often kept off the "critical path" of decision-making because the latency made real-time reasoning impractical. At 750 tokens per second, that constraint is gone. AI can now be embedded in incident response, financial trading, customer support escalations, and real-time commerce—workflows where every second counts.

OpenAI Just Broke the AI Triangle. Speed and Intelligence No Longer Trade Off.
OpenAI's Ultrafast mode runs GPT-5.6 Sol at 14× the speed without sacrificing intelligence.

The $10 Billion Cerebras Bet Just Paid Off

Ultrafast also has implications for the broader chip war. Nvidia has reportedly spent $20 billion on Groq's SRAM-based technology to achieve similar low-latency inference. AMD, meanwhile, has partnered with Cerebras on disaggregated inference for its Helios racks, using GPUs for prompt processing and Cerebras for token generation. OpenAI's Ultrafast is a validation of Cerebras' architectural bet—and a signal that the industry is moving beyond GPUs for inference.


P.S. The preview is currently limited. OpenAI is starting with a small group of customers to "learn where that speed creates meaningful value" before expanding the service. Capacity is the constraint—Cerebras' hardware does not yet scale to full production demand. But the direction is clear: frontier intelligence is no longer just about capability. It is about how fast it can deliver.


Frequently Asked Questions

Q: What is OpenAI's Ultrafast mode?

A: Ultrafast is a new service tier for GPT-5.6 Sol that delivers up to 750 output tokens per second—approximately 14× faster than Standard processing—without sacrificing intelligence or accuracy.

Q: How is Ultrafast able to run so fast?

A: Ultrafast is powered by Cerebras' Wafer-Scale Engine architecture, which keeps model weights entirely on-chip in SRAM, eliminating the memory bandwidth bottleneck that constrains traditional GPU inference. This allows for near-instantaneous token generation.

Q: What are the benchmark results for Ultrafast?

A: On Humanity's Last Exam (2,500 graduate-level questions), Ultrafast completed the set in 11 hours vs 78 hours for Fable 5. On GDP-Val, it delivered a 5.6× end-to-end speedup. It's also 5× faster than Opus 4.8 Fast and 11× faster than Fable 5.

Q: What are the early use cases for Ultrafast?

A: OpenAI's internal teams are using it for incident response (analyzing logs during active outages), research iteration (running experiments multiple times per day), and other workflows where real-time intelligence is critical. Early external customers include Jane Street, Basis, Rogo, and Podium.

Q: Is Ultrafast available to everyone?

A: No. It's in limited preview to a small group of customers. OpenAI is starting with a small group to "learn where that speed creates meaningful value" before expanding.

Q: What is Cerebras' role in this?

A: Cerebras builds Wafer-Scale Engine chips that keep 44GB of model weights entirely on-chip in SRAM. OpenAI signed a $10+ billion deal with Cerebras in January 2026 to deploy 750 megawatts of accelerators through 2028. Ultrafast is the first product of that partnership.

Q: Does Ultrafast sacrifice accuracy for speed?

A: No. OpenAI emphasizes that Ultrafast delivers "frontier intelligence" with no loss in accuracy. The speed gain comes from the hardware architecture, not from model compression or simplification.

Q: How does this compare to Nvidia and AMD?

A: Nvidia has reportedly spent $20 billion on Groq's SRAM-based technology for similar low-latency inference. AMD has partnered with Cerebras on disaggregated inference for its Helios racks. OpenAI's Ultrafast validates Cerebras' architectural approach and signals that the industry is moving beyond GPUs for inference.

Q: What is the pricing for Ultrafast?

A: OpenAI has not released specific pricing details for the Ultrafast tier. The service is currently in preview with limited customer access.

Q: When will Ultrafast be widely available?

A: OpenAI has not announced a timeline for general availability. The preview phase is designed to "learn where that speed creates meaningful value" before expanding capacity and access.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article