OpenAI's Jalapeño: RTL freeze to tapeout in 9 months, serving ChatGPT 10 weeks after first silicon

thehiphopswami · x · 2026-09-04

The Real Story of OpenAI's Jalapeño at Hot Chips

The specs are solid but not the headline: 6 HBM4 stacks, 216GB at 15.4 TB/s, 700W, designed around speculative decoding. OpenAI claims 1.7x tokens per kW over GB300 on DeepSeek R1 (power-normalized) — while handicapping themselves, letting the GPU baseline use speculative decoding while Jalapeño ran plain single-token prediction.

The Process Is the Story

The author argues this chips away at a 40-year moat: AI-driven chip design is itself the new methodology.

Original post →

More from Infra

Infra channel →