OpenAI's Jalapeño inference chip: 9 months to tapeout, one chip beats GPU+LPU pair
thehiphopswami · x · 2026-08-28
A new SemiAnalysis podcast episode takes a deep dive into OpenAI's Jalapeño inference chip announced at Hot Chips:
- "Dark silicon is cheaper than idle accelerators": one balanced chip may beat a GPU + LPU pair
- Nine months from RTL to tapeout, a Blackwell-class chip designed with AI for EDA
- NUMA-style local HBM slices per accelerator fix the "operands arrive late" problem
- Broadcom ESUN scale-up: 128 chips per rack at 600 GB/s, 2,048 chips across 16 racks still count as scale-up
- The "regret factor": the opportunity cost of missing a future model beats the cost of generality
Plus spicy extras like whether OpenAI should sell the chip externally.
More from Infra
- Miles now supports RL training for Qwen and GLM models with high-performance kernels — ying11231 · 2026-08-28
- AI Semiconductor Endgame 2026: Infra Bubble and Open vs Closed Economics — sudoraohacker · 2026-08-28
- GLM-5.3-Flash hits 270 tok/s: 10% higher quality than 5.2 at one-tenth the cost — Yuchenj_UW · 2026-08-28
- Ornith-1.5-35B-A3B runs agentic coding at 32 tok/s on an 8GB RTX 3070 laptop — Elemental_Particle · 2026-08-28
- From MySQL to Redis: Internet Scaling History and AI Lessons — generativist · 2026-08-28
- Puro-2B Matches Qwen2.5 Performance with $6.9K Pretraining on RTX 5090 — _reachsumit · 2026-08-28