OpenAI's first in-house chip Jalapeño claims wins over Nvidia flagships in inference benchmarks
OpenAI has released benchmark results for Jalapeño, its first in-house inference chip, claiming it beats NVIDIA's flagship GB200 and GB300 systems across multiple inference workloads. The chip was co-developed with Broadcom, going from team formation to tape-out in only about 16 months. The conclusions currently come from OpenAI's own InferenceX tests and an independent teardown analysis by SemiAnalysis, and are drawing wide attention because they could shake NVIDIA's position in the AI inference compute market.
Confirmed
- OpenAI unveiled Jalapeño, its first self-developed inference chip, co-developed with Broadcom, with a design cycle of only about 16 months (@dylan522p, @zephyrz9).
- On OpenAI's public InferenceX benchmark, tested with open-source models including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2, Jalapeño delivers higher peak throughput per kilowatt and lower token latency, with better TCO and throughput than NVIDIA (@m2, @m4, @m6).- At the same DeepSeek R1 decode speed, throughput per kilowatt is 104.3x that of NVIDIA GB300 (12,258 vs 118 tok/s/kW) (@rohanpaulai).
- Energy efficiency reaches up to 1.9x that of Nvidia GB200/GB300 (@kimmonismus).
- SemiAnalysis's deep-dive teardown notes that Jalapeño is not narrowly optimized only for OpenAI models but is a general-purpose LLM inference ASIC, with measured efficiency exceeding NVIDIA and AMD (@zephyrz9, @dylan522p).
- OpenAI says that versus NVIDIA Blackwell, Jalapeño leads on two categories of metrics including AI processing per unit of power (official statements relayed by @dinabass).
- The chip is positioned to improve model inference efficiency and lower compute costs (@Polymarket).
Not yet confirmed
- All performance figures come from OpenAI's own InferenceX benchmark and SemiAnalysis's analysis; there is no independent verification from NVIDIA or any authoritative third party yet. The test conditions behind extreme numbers like the "104.3x" claim (e.g., comparison system configurations, quantization, and batching settings) still await more public detail.
Why it matters
- This is OpenAI's first public disclosure of real-world data for its in-house chip, marking the landing phase of its strategy to reduce dependence on NVIDIA. If the general-purpose inference ASIC positioning holds up, it could put real competitive pressure on NVIDIA and AMD in the data center inference market.
2026-08-25 ~ 2026-08-25 · 12 related posts
Primary sources
- OpenAI's first custom inference chip Jalapeño claims 1.5-1.9x better perf-per-watt than Nvidia GB200/GB300 — kimmonismus ·
- OpenAI's 'Jalapeño' Chip Beats Nvidia Blackwell in Benchmarks — dylan522p ·
- SemiAnalysis: OpenAI's Jalapeño Is a General-Purpose Inference ASIC Beating NVIDIA and AMD in Tests — zephyr_z9 ·
- [source] SemiAnalysis: OpenAI's Jalapeño Is a General-Purpose Inference ASIC Beating NVIDIA and AMD in Tests — zephyr_z9 · 2026-08-25
- OpenAI Reveals Jalapeño Benchmarks: Beats Blackwell in Efficiency — dinabass · 2026-08-25
- Media Reports on OpenAI Jalapeño Benchmark Data — shiringhaffary · 2026-08-25
- [source] OpenAI's 'Jalapeño' Chip Beats Nvidia Blackwell in Benchmarks — dylan522p · 2026-08-25
- OpenAI claims custom 'Jalapeño' chip outperforms Nvidia GB200/GB300 in inference — Polymarket · 2026-08-25
- OpenAI's Jalapeño chip claims 104.3x better efficiency than Nvidia GB300 — rohanpaul_ai · 2026-08-25
- [source] OpenAI's first custom inference chip Jalapeño claims 1.5-1.9x better perf-per-watt than Nvidia GB200/GB300 — kimmonismus · 2026-08-25
- OpenAI Claims New 'Jalapeño' Chip Outperforms Vera Rubin in Benchmarks — Wonderful_Buffalo_32 · 2026-08-25
- Jalapeño Claims Industry-Leading AI Inference Speed and Efficiency in First Results — pstAsiatech · 2026-08-25
3 near-duplicate retellings: firstadopter · thesaraharminta · firstadopter