OpenAI's Jalapeño: RTL freeze to tapeout in 9 months, serving ChatGPT 10 weeks after first silicon
thehiphopswami · x · 2026-09-04
The Real Story of OpenAI's Jalapeño at Hot Chips
The specs are solid but not the headline: 6 HBM4 stacks, 216GB at 15.4 TB/s, 700W, designed around speculative decoding. OpenAI claims 1.7x tokens per kW over GB300 on DeepSeek R1 (power-normalized) — while handicapping themselves, letting the GPU baseline use speculative decoding while Jalapeño ran plain single-token prediction.
The Process Is the Story
- A tiny team with barely any traditional chip design experience
- A new Rust-like hardware description language built so AI writes correct code out of the box
- An internal AI optimization loop took the design from functionally correct to beyond expert level
- RTL freeze to tapeout in 9 months; first silicon to serving ChatGPT in 10 weeks
- Gen 2 already approaching tapeout
The author argues this chips away at a 40-year moat: AI-driven chip design is itself the new methodology.
More from Infra
- ASML relies heavily on FPGAs, and Xilinx once caused major backlogs — PAstynome · 2026-09-04
- Tenstorrent Launches JapanFold: Sovereign Open-Source Drug Discovery for Japan — DavidBennett__ · 2026-09-04
- Stanford's Prefix Sliding cuts long-reasoning inference ~3x without retraining, AIME25 score intact — rohanpaul_ai · 2026-09-04
- Analyst: no model trained without NVIDIA has ever beaten one trained on NVIDIA — BenBajarin · 2026-09-04
- A 2023 desktop PC now resells above its purchase price amid hardware inflation — Darpinian · 2026-09-04
- Lenovo Unveils Dozens of AI Devices at IFA, RTX Spark Laptop Runs 120B-Param Models Locally — 智东西 · 2026-09-04