OpenAI's Jalapeño Chip Benchmarks: 1.9x Efficiency Over Nvidia GB200

AI寒武纪 · wechat · 2026-08-26

OpenAI revealed benchmarks for its first in-house inference chip, Jalapeño. Tested on SemiAnalysis's InferenceX, Jalapeño delivers 1.5x to 1.9x more tasks per watt and 1.7x to 3.6x lower end-to-end latency compared to Nvidia GB200/GB300, with performance in high-interaction scenarios reaching 2.1x to 4.1x. The chip runs at 550W or lower under load despite a 700W rating. Its architecture features full-stack hardware-software co-design to minimize data movement and explicitly manage KV cache, optimizing for agent workloads. The 9-month development cycle leveraged AI for design and verification, with AI-generated code running 1.8x faster than human code in some cases. OpenAI plans deployment by year-end, with Gen 2 and Gen 3 chips already in development.

Related event: OpenAI's First Custom Inference Chip Jalapeño Beats Nvidia Flagships(45 posts)→

Original post →

More from Infra

Infra channel →