Story
August 25, 2026
OpenAI’s Jalapeño Chip Takes Aim at Nvidia, but Dependence Isn’t Over
OpenAI says its first custom inference chip delivers major gains in speed and energy efficiency. Yet Jalapeño does not train models, deployment remains limited at first, and Nvidia will stay central to OpenAI’s compute plans.
OpenAI’s Jalapeño is a pointed attempt to loosen Nvidia’s grip on the AI boom: a chip built to make model responses faster and cheaper. But the company’s promising benchmark debut is not yet a declaration of hardware independence.
First announced last October and developed with Broadcom, Jalapeño was presented Tuesday at Hot Chips as OpenAI’s first custom inference processor. Tests on SemiAnalysis’ InferenceX benchmark put it ahead of available systems on throughput per kilowatt and response speed, including a Nvidia Blackwell comparison. Richard Ho, OpenAI’s hardware chief, called the results “a very, very significant performance advance over state of the art.”1
OpenAI’s own account makes the case more broadly: across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T, it says Jalapeño delivered 1.5 to 1.9 times more work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparison systems.2 The claimed edge rests on tailoring the chip, memory, networking and serving software to the distinct prefill and decode stages of language-model inference, reducing data movement and delays.
That full-stack approach is the strategic prize. OpenAI says owning more of the infrastructure gives it greater control over the economics of serving models, while preserving a broad supplier portfolio that includes Microsoft, Nvidia, AMD, AWS, Broadcom and others.3 In other words, Jalapeño is leverage—not a wholesale replacement plan.
The limits are clear. The chip is designed for inference, not the training of frontier models, leaving OpenAI reliant on Nvidia and other suppliers for the immense compute buildout ahead. “We’re going to need a lot of compute,” Ho said. “And Jalapeño is part of that. Cerberus is part of that. Nvidia is part of that. AMD is part of that.”4
OpenAI says deployment will begin in limited form before scaling through later generations. The decisive test will be whether those lab results survive production at scale—while rival chips keep advancing.