
OpenAI
AI Hardware · Semiconductors · Cloud Infrastructure
OpenAI's Jalapeño chip beats Nvidia Blackwell on AI inference speed
August 25, 2026
Reducing dependence on a single GPU supplier could reshape the economics of running AI models at scale for every major lab.
- OpenAI presented the first benchmark results for Jalapeño, its custom inference chip built with Broadcom, at the Hot Chips 2026 conference on August 25.
- Jalapeño is an inference-only ASIC designed to run large language models behind ChatGPT, Codex, and OpenAI's API rather than to train new models.
- Across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5, Jalapeño delivered 1.5-1.9x more AI work per watt and 1.7-3.6x lower end-to-end latency than the comparison Nvidia systems.
- Tests ran on SemiAnalysis's power-normalized InferenceX benchmark, with Jalapeño operating at 700 watts versus 1.2-1.4 kilowatts for the Nvidia GB200/GB300 and AMD MI355X systems it was measured against.
- OpenAI plans to deploy Jalapeño in small volumes by the end of 2026, ramping to larger-scale production in 2027, while already developing its second and third generation chips.
- OpenAI still relies on Nvidia and AMD GPUs for training and near-term inference, joining Google, Amazon, and Meta in building custom AI silicon alongside, not instead of, its GPU suppliers.
- Inference, not training, is now the fastest-growing and most margin-sensitive segment of AI compute, so a competitive in-house chip threatens Nvidia's pricing power exactly where usage scales fastest.