tech
Jalapeño's first results show industry-leading speed and efficiency in AI inference
Since announcing Jalapeño, OpenAI’s first custom inference chip, we have been testing the chip and the system built around it. The results show a significant performance advance: Jalapeño can serve more AI work per unit of power while also returning responses more quickly. Jalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two.

TL;DR
- Jalapeño, OpenAI's first custom inference chip, shows significant performance gains in speed and power efficiency.
- It offers both higher throughput and lower latency with a single architecture, unlike systems that typically trade one for the other.
- Performance improvements were observed across various AI models, including GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T.
- Jalapeño delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency compared to benchmark systems.
- The chip's development was accelerated by OpenAI's own AI models, showcasing a full-stack advantage in designing hardware and software together.
- OpenAI plans to deploy Jalapeño within its compute infrastructure by the end of the year, with future generations already in development.