OpenAI's Jalapeño Chip Benchmarks: 1.9x Efficiency Over Nvidia GB200
AI寒武纪 · wechat · 2026-08-26
OpenAI revealed benchmarks for its first in-house inference chip, Jalapeño. Tested on SemiAnalysis's InferenceX, Jalapeño delivers 1.5x to 1.9x more tasks per watt and 1.7x to 3.6x lower end-to-end latency compared to Nvidia GB200/GB300, with performance in high-interaction scenarios reaching 2.1x to 4.1x. The chip runs at 550W or lower under load despite a 700W rating. Its architecture features full-stack hardware-software co-design to minimize data movement and explicitly manage KV cache, optimizing for agent workloads. The 9-month development cycle leveraged AI for design and verification, with AI-generated code running 1.8x faster than human code in some cases. OpenAI plans deployment by year-end, with Gen 2 and Gen 3 chips already in development.
Related event: OpenAI's First Custom Inference Chip Jalapeño Beats Nvidia Flagships(45 posts)→
More from Infra
- Lighter crypto system achieves >10x performance boost via optimization — gajesh · 2026-08-26
- Antirez on Mac Studio M5 Ultra: Compute/bandwidth insufficient for K3 tasks — antirez · 2026-08-26
- Evaluating Mac Studio M5 Ultra value by GLM 5.3 concurrency and speed — antirez · 2026-08-26
- EVE Online begins migration: upgrading 2.4M lines of code to Python 3 — Simon Willison · 2026-08-26
- IBM Granite 4.2 8b Blooms to 27GB VRAM Despite 5.2GB Download — x8code · 2026-08-26
- Hardcore Player Builds Underground Data Center Just to Heat Outdoor Jacuzzi — curious_vii · 2026-08-26