SemiAnalysis: OpenAI's Jalapeño Is a General-Purpose Inference ASIC Beating NVIDIA and AMD in Tests
zephyr_z9 · x · 2026-08-25
SemiAnalysis shares more details on OpenAI's Jalapeño chip:
- Not narrowly optimized for OpenAI models — it's a general-purpose LLM inference ASIC, designed from scratch, going from team buildout to tape-out in roughly 16 months.
- In their tested workloads it outperformed NVIDIA, AMD and Google accelerators on several inference efficiency metrics.
- Performance: DeepSeek R1 exceeded 700 tokens/sec/user at concurrency 1; Kimi K2.5 and GPT-OSS reached roughly 1,400 tokens/sec/user in selected tests.
- These results used single-token prediction, without speculative decoding or prefill/decode disaggregation.
More from Infra
- Australia Faces Datacentre Rush as AI Firms Scramble to Dodge Upcoming Regulations — nordicinst · 2026-08-25
- Jalapeño ASIC outperforms comparable TPUs, challenging existing giants — GavinSBaker · 2026-08-25
- Beyond model quality: cheaper, faster inference may decide the AI race — why OpenAI's full-stack bet matters — VraserX · 2026-08-25
- Samsung, SK Hynix step up China NAND upgrades as US tool risks grow — pstAsiatech · 2026-08-25
- Samsung Raises Advanced Chip Prices by Up to 15% Amid Surging Demand — Beth_Kindig · 2026-08-25
- China Pioneers 'Know Your Agent' Rules for AI Payments — pstAsiatech · 2026-08-25