
OpenAI
AI Hardware · Semiconductors · AI Infrastructure
OpenAI's Jalapeño chip beats Nvidia GB300 in first benchmarks
August 25, 2026
Building inference silicon in-house lets OpenAI tune chips, models, and software together instead of waiting on suppliers' roadmaps.
- OpenAI published its first independent benchmark results for Jalapeño, a custom inference chip built with Broadcom, at the Hot Chips conference on August 25, 2026.
- Jalapeño only handles inference (running trained models), not training, and OpenAI says it's a general-purpose accelerator rather than one tuned exclusively to its own models.
- Tested on SemiAnalysis's InferenceX benchmark, Jalapeño delivered 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower end-to-end latency than comparison Nvidia GB200/GB300 systems across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.
- The 700W Jalapeño chip outperformed Nvidia's 1,400W GB300 flagship on throughput-per-watt, and for interactive agent workloads posted up to 4.1x higher performance.
- OpenAI plans to deploy Jalapeño in small volumes by the end of 2026 and ramp meaningfully in 2027, with second- and third-generation chips already in development.
- Jalapeño was first unveiled in June 2026 as an OpenAI-Broadcom partnership; this Hot Chips presentation is the first time OpenAI has released concrete performance numbers.
- Owning inference silicon gives OpenAI leverage against Nvidia's pricing and supply constraints as it races to serve growing model demand more cheaply and reliably.