OpenAI revealed its first self-developed AI inference chip, Jalapeño, at the Hot Chips conference on August 25, demonstrating performance gains over Nvidia's GB200 and GB300 GPUs in power efficiency and latency metrics [1, 2, 3, 4, 5, 6]. The chip delivered 1.5 to 1.9 times higher AI throughput per watt at peak load and reduced end-to-end latency by 1.7 to 3.6 times across three open-source AI models, according to benchmarking run by SemiAnalysis in OpenAI labs [1, 2, 3, 4, 5, 6]. Richard Ho, OpenAI's hardware lead, said, "Jalapeño can simultaneously provide high throughput and low latency, optimizing power usage to lower data center electricity and operational costs" [2].
The Jalapeño chip has a thermal design power (TDP) of approximately 700 watts, considerably lower than Nvidia's GB200 (1,200 watts) and GB300 (1,400 watts) GPUs. It also improves on the TDP projected for Nvidia's upcoming Vera Rubin chip, expected between 900 and 1,150 watts [1, 3, 6]. Jensen Huang, Nvidia CEO, noted that "throughput per watt is revenues for data center operators with fixed power allocation," highlighting the importance of the efficiency gains [1].
Focused on AI inference rather than training, Jalapeño is optimized primarily for large language model workloads including GPT-OSS 120B, DeepSeek-R1, and Moonshot AI models [1, 2, 5, 6]. Its theoretical compute capacity reaches about 13.4 PFLOPs, supported by 15.4 TB/s memory bandwidth [6]. The chip was developed in roughly 16 months starting mid-2024, with Broadcom's collaboration and fabrication using TSMC’s N3P process and HBM4 memory technology [2, 3, 4, 6]. Ho explained the architecture was tuned to match advanced AI model characteristics, leveraging OpenAI’s own Codex AI and a new GPU programming language called Gluon to approach hardware theoretical limits [3, 4, 6].
The benchmarking conducted mainly covered 8k1k workloads, with full benchmark suite results still pending [6]. One reported metric found Jalapeño achieving up to 104 times faster token decoding speeds in high concurrency tests compared to Nvidia GB300 in a specific scenario [4]. However, discussions remain about the completeness of comparisons to Nvidia's newest Blackwell architecture and Vera Rubin chip, which were not fully tested in these early results [1, 2, 3, 5, 6]. OpenAI clarified that Jalapeño is intended to complement rather than replace Nvidia products, stressing continued partnership with Nvidia for various chips and platforms [2, 3, 5].
OpenAI's Jalapeño chips are presently assembled by Celestica, with Taiwanese ODM firms like Foxconn, Quanta, Wistron, Welyt, and Inventec anticipated to support future scaling [4, 6]. The company plans limited deployment of Jalapeño in its AI services by the end of 2026. Mass production ramp-up is scheduled for 2027 [2, 3, 5, 6].