OpenAI's Jalapeño ASIC beats NVIDIA GB300 at half the power: 1.5-1.9x throughput per watt
新智元 · wechat · 2026-09-10
Anthropic has confirmed an internal custom silicon team for Claude inference chips, while OpenAI unveiled full benchmarks for its Jalapeño inference ASIC at HotChips: 700W vs NVIDIA's 1200W GB200 and 1400W GB300, with 1.5-1.9x higher throughput per watt and 1.7-3.6x lower end-to-end latency. The same week, Qualcomm signed a $60B custom inference chip deal with Amazon.
- Agents reshape compute: inference shifts from one-shot requests to long-running task chains (multi-turn reasoning, tool calls, retries), demanding always-on, recoverable, cost-controlled systems.
- Jalapeño's three-stage design: Prefill (compute-heavy), Draft (speculative sampling, latency-sensitive), Verification (bandwidth-heavy with bursty MoE traffic) on one chip, local KV cache, powering down idle units — the key to beating 1400W parts at 700W.
- China's gap: rich applications and diverse chip routes (RISC-V, Chiplet, compute-in-memory) but high interop/migration costs risk fragmentation.
- The piece ends promoting the AICC2026 conference in Beijing and a China AI compute report.
More from Infra
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11