tech

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

TL;DR

  • OpenAI's Jalapeño chip shows significant performance gains in benchmark tests.
  • Jalapeño outperforms current state-of-the-art inference processors in tokens per user and throughput per kilowatt.
  • The chip was developed by OpenAI in collaboration with Broadcom, with AI models assisting in the process.
  • A full-stack design approach minimizes data movement and communication delays during inference.
  • Jalapeño is intended to be a multigenerational platform.
  • Limited deployment is expected by the end of 2026, with wider deployment in 2027.