Hardware

OpenAI's Jalapeño Chip Just Shattered the Inference Trade-Off

CRAZE CRAZE Summary 3 things to know
  • OpenAI's custom Jalapeño chip beats Nvidia's GB200/GB300 by 1.5–1.9x per watt and up to 4x lower latency.
  • Built with Broadcom on TSMC 3nm, it went from design to tapeout in nine months using AI-assisted chip design.
  • OpenAI won't sell Jalapeño; it's a vertically integrated inference play, with gen two in development and 2027 deployment.
Jeff Editorial | · 4 min read
OpenAI's Jalapeño Chip Just Shattered the Inference Trade-Off

At the Hot Chips conference on August 25, OpenAI's VP of hardware Richard Ho presented the first performance data for Jalapeño, the company's custom inference ASIC developed in partnership with Broadcom and manufactured on TSMC's 3nm process. The chip is rated at 700W, though measured sustained power remained at or below 550W during testing. A rack of 128 chips delivers 1.7 exaflops of 4-bit compute with 27.5TB of HBM4 memory.

OpenAI's Jalapeño Chip Just Shattered the Inference Trade-Off
OpenAI's custom inference chip Jalapeño delivers 1.9x more AI work per watt and 4x lower latency than Nvidia's best.

1.9x Per Watt, 4x Faster—No Trade-Off

OpenAI tested Jalapeño against Nvidia's GB200 and GB300 systems using SemiAnalysis' InferenceX benchmark, running three models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T.

Metric

Jalapeño vs. GB200/GB300

AI work per watt

1.5x – 1.9x higher

End-to-end latency reduction

1.7x – 3.6x lower

Ultra-low latency performance

2.1x – 4.1x faster

On GPT-OSS 120B, Jalapeño delivered roughly 1.9x higher peak throughput per kilowatt than GB200. On DeepSeek R1, it outperformed GB300 by 1.7x per watt; on Kimi K2.5, by 1.5x. The performance advantage widened further on OpenAI's internal frontier models, suggesting the architecture becomes more effective as workloads grow larger.

OpenAI's Jalapeño Chip Just Shattered the Inference Trade-Off
Jalapeño delivers 1.9x higher throughput per kilowatt than Nvidia's GB200 on GPT-OSS 120B, with up to 4x lower latency.

From Design to Tapeout in 9 Months

Jalapeño went from initial design to tapeout in nine months—one of the fastest ASIC development cycles in the history of high-performance semiconductors. OpenAI credits its own AI models for accelerating the process.

Using Codex with GPT-Astra, the team brought three open-weight models to high performance within two months. For selected GPT-OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than existing human-written implementations.

OpenAI's hardware VP Richard Ho described the chip as a "general purpose, very flexible accelerator," noting that it worked across internal and external models. Jalapeño is not designed for training—only inference—meaning OpenAI remains dependent on Nvidia and others for training workloads.

OpenAI Is Now a Chip Company. They're Not Selling.

OpenAI has no plans to sell Jalapeño to other companies. "We have so much need for it that we can't imagine when we would be able to," Ho said. The chip is part of a broader compute strategy that includes Nvidia, Cerebras, and AMD.

The implications are clear: the largest AI model developers are no longer just customers of chip companies—they are becoming chip companies. Google has TPUs. Amazon and Microsoft have their own designs. Anthropic confirmed earlier this month it is looking to build its own in-house chip design effort.

Jalapeño is expected to be deployed in limited quantities by the end of 2026, with wider deployment in 2027. A second-generation chip is already "deep into development," and a third generation has begun concept design.


P.S. The chip's advantage on internal frontier models—not just open-source benchmarks—suggests OpenAI designed Jalapeño to optimize for its own architecture as much as for general inference. The company is building a vertically integrated AI stack, from model to silicon. And it just proved it works.


Frequently Asked Questions

Q: What is OpenAI's Jalapeño chip?

A: Jalapeño is OpenAI's first custom inference ASIC, developed in partnership with Broadcom and manufactured on TSMC's 3nm process. It is designed to run AI models efficiently during inference, not training, and is optimized for OpenAI's own model architectures.

Q: When was Jalapeño announced?

A: The first performance data for Jalapeño was presented by OpenAI's VP of hardware Richard Ho at the Hot Chips conference on August 25, 2026.

Q: How does Jalapeño compare to Nvidia's best chips?

A: On GPT-OSS 120B, Jalapeño delivered roughly 1.9x higher peak throughput per kilowatt than Nvidia's GB200, with up to 4x lower latency. On DeepSeek R1, it outperformed GB300 by 1.7x per watt; on Kimi K2.5, by 1.5x. The performance advantage widened further on OpenAI's internal frontier models.

Q: What are the key specs?

A: A rack of 128 chips delivers 1.7 exaflops of 4-bit compute with 27.5TB of HBM4 memory. The chip is rated at 700W, though measured sustained power remained at or below 550W during testing.

Q: What is the "trade-off" that Jalapeño eliminates?

A: Traditional inference systems force a choice between high throughput (processing many requests) and low latency (responding quickly). Jalapeño delivers both—higher throughput and lower latency—in a single architecture without sacrificing efficiency.

Q: Is Jalapeño available to buy?

A: No. OpenAI has no plans to sell Jalapeño. The chip is designed for OpenAI's own use. As Richard Ho said: "We have so much need for it that we can't imagine when we would be able to."

Q: How fast was the development cycle?

A: Jalapeño went from initial design to tapeout in nine months—one of the fastest ASIC development cycles in the history of high-performance semiconductors. OpenAI credits its own AI models for accelerating the process.

Q: What is the timeline for deployment?

A: Jalapeño is expected to be deployed in limited quantities by the end of 2026, with wider deployment in 2027. A second-generation chip is already "deep into development," and a third generation has begun concept design.

Q: Why does OpenAI need its own chip?

A: Jalapeño is part of a broader compute strategy that includes Nvidia, Cerebras, and AMD. By building its own inference chip, OpenAI can optimize for its own model architectures, reduce reliance on third-party suppliers, and improve inference efficiency at scale.

Q: Does this mean OpenAI is becoming a chip company?

A: Yes—and they are not alone. Google has TPUs. Amazon and Microsoft have their own designs. Anthropic confirmed earlier this month it is looking to build its own in-house chip design effort. The largest AI model developers are no longer just customers of chip companies. They are becoming chip companies.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article