FULL STORY

OpenAI's Jalapeño Chip: Benchmarks, Deployment and the Road Ahead

OpenAI unveiled benchmark results for its in-house Jalapeño inference chip, claiming wins over Nvidia's GB200/GB300, with deployment planned by year-end and multiple generations already in the pipeline.

2026-08-25 ~ 2026-08-26 · 3 episodes · 52 posts

Episode 1 · OpenAI's First Custom Inference Chip Jalapeño Beats Nvidia in Tests (2026-08-25, 48 posts)

OpenAI has released benchmark results for Jalapeño, its first in-house inference chip, claiming it beats NVIDIA's flagship GB200 and GB300 systems across multiple inference workloads. The chip was co-developed with Broadcom, going from team formation to tape-out in only about 16 months. The conclusions currently come from OpenAI's own InferenceX tests and an independent teardown analysis by SemiAnalysis, and are drawing wide attention because they could shake NVIDIA's position in the AI inference compute market.

Confirmed

  • OpenAI unveiled Jalapeño, its first self-developed inference chip, co-developed with Broadcom, with a design cycle of only about 16 months (@dylan522p, @zephyrz9).
  • On OpenAI's public InferenceX benchmark, tested with open-source models including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2, Jalapeño delivers higher peak throughput per kilowatt and lower token latency, with better TCO and throughput than NVIDIA (@m2, @m4, @m6).- At the same DeepSeek R1 decode speed, throughput per kilowatt is 104.3x that of NVIDIA GB300 (12,258 vs 118 tok/s/kW) (@rohanpaulai).
  • Energy efficiency reaches up to 1.9x that of Nvidia GB200/GB300 (@kimmonismus).
  • SemiAnalysis's deep-dive teardown notes that Jalapeño is not narrowly optimized only for OpenAI models but is a general-purpose LLM inference ASIC, with measured efficiency exceeding NVIDIA and AMD (@zephyrz9, @dylan522p).
  • OpenAI says that versus NVIDIA Blackwell, Jalapeño leads on two categories of metrics including AI processing per unit of power (official statements relayed by @dinabass).
  • The chip is positioned to improve model inference efficiency and lower compute costs (@Polymarket).

Not yet confirmed

  • All performance figures come from OpenAI's own InferenceX benchmark and SemiAnalysis's analysis; there is no independent verification from NVIDIA or any authoritative third party yet. The test conditions behind extreme numbers like the "104.3x" claim (e.g., comparison system configurations, quantization, and batching settings) still await more public detail.

Why it matters

  • This is OpenAI's first public disclosure of real-world data for its in-house chip, marking the landing phase of its strategy to reduce dependence on NVIDIA. If the general-purpose inference ASIC positioning holds up, it could put real competitive pressure on NVIDIA and AMD in the data center inference market.

28 more related posts →

Episode 2 · OpenAI to Deploy Its Jalapeño Chip by Year-End with Next Generations in Development (2026-08-26, 2 posts)

OpenAI plans to deploy its custom Jalapeño chip by year-end as the first step in a multi-generation roadmap, with unconfirmed reports of major internal performance gains. The second-generation chip is in deep development and a third is taking shape.

Episode 3 · OpenAI's Jalapeño Chip Play: Winning Either Way (2026-08-26, 2 posts)

SemiAnalysis analysis of OpenAI's custom ASIC, codenamed Jalapeño, suggests OpenAI wins regardless of deployment outcome, potentially extracting billions from Nvidia deals. Jalapeño reportedly outperforms Cerebras on cost and power efficiency.