OpenAI's Custom Inference Chip Jalapeño Delivers 1.9x More Work Per Watt Than Nvidia GB300

ivan_bezdomny · x · 2026-08-26

OpenAI's first custom inference chip 'Jalapeño' reportedly delivers 1.5-1.9x more work per watt and 1.7-3.6x lower latency than Nvidia's GB300 system. Built on TSMC N3P with 15.4TB/s bandwidth, the 700W chip was designed using AI assistance and is scheduled for deployment by late 2026.

Related event: OpenAI Unveils First Custom Inference Chip Jalapeño with Benchmark Results(42 posts)→

Original post →

More from Infra

Infra channel →