OpenAI released test results on August 26 showing its new Jalapeño AI inference processor surpasses Nvidia’s leading GPUs in speed and energy efficiency. [1, 2, 3] Developed with Broadcom, the chip delivers superior AI workload performance per watt, operating efficiently at a thermal design power of around 700 watts, which reduces data center costs, according to OpenAI’s chip chief Richard Ho. [1, 2, 3]

Jalapeño achieved higher token throughput and faster response times compared to Nvidia’s GB300 and Blackwell chips. It also outperformed Nvidia’s upcoming Rubin generation in early engineering tests on energy efficiency and token throughput, although industry analysts noted Nvidia remains dominant for broader AI compute tasks beyond inference. [1, 2, 3] Jalapeño uses TSMC’s advanced N3P semiconductor process, reaching a theoretical peak of 13.4 petaflops and memory bandwidth of 15.4 terabytes per second with HBM4 memory. [3]

Ho described the chip as excelling in both high-throughput and low-latency AI workloads, saying, “It’s a really good chip — it should drop [costs] by a lot.” Adrien Sanchez, an analyst at Yole Group, commented, “A hyperscaler-designed chip can now match or beat Nvidia's Blackwell-class GPUs on inference efficiency.” [1, 2]

OpenAI began designing Jalapeño in mid-2024 and completed tape-out after about 16 months. They used their AI coding tool Codex and a custom language named Gluon for programming the chip and related software. [3] Second and third generations of the chip are already in development. [1, 2]

OpenAI plans to deploy Jalapeño chips in its compute infrastructure by the end of 2026, aiming to reduce reliance on Nvidia GPUs for AI inference tasks over time. [1, 2] Full production scaling is expected in 2027. [3]

While Nvidia still leads in customizable GPUs suited for broader workloads, OpenAI’s Jalapeño highlights a growing trend of companies designing specialized AI chips to optimize inference efficiency. The rollout of Jalapeño later this year marks a significant step in OpenAI’s hardware capabilities.