OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput
firstadopter · x · 2026-09-01
Tae Kim reports on OpenAI's custom chip 'Jalapeno' presented at Hot Chips. Key details include:
- Full-stack optimization: Completed RTL execution in nine months, optimizing from models down to silicon.
- Performance: Achieved nearly 1500 tokens/s on OSS models with 4x lower latency, and 700 tokens/s on DeepSeek with 5x lower latency.
- Design focus: Dedicated to accelerating OpenAI workloads to reduce user wait times and energy consumption.
More from Infra
- Google Cloud Monitoring MCP Connector Released — modelcontextprotocol · 2026-09-01
- Distributed.systems发布可审计的Agent基础设施 — arthurcolle · 2026-09-01
- Does enabling ChatGPT Memory or history reference increase token usage? — ssunki · 2026-09-01
- Engineer fixes ROCm inference crash on MI350X, uncovers 9 bugs in deep dive — AnushElangovan · 2026-09-01
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01