OpenAI's First Custom Inference Chip Jalapeño Beats Commercial Systems
firstadopter · x · 2026-08-25
OpenAI released the first measured performance results for Jalapeño, its first custom inference chip:
Performance:
- Delivers higher peak throughput per kilowatt and lower token latency than commercial systems on the InferenceX benchmark (GPT-OSS 120B).
- Shows strong performance on DeepSeek R1 and Kimi K2, extending gains across model families.
Development & Efficiency:
- AI-Assisted Design: AI enabled the team to move from initial design to tapeout in nine months, shortening verification loops.
- Circuit Optimization: AI helped optimize arithmetic circuits to fit more compute performance.
- Jevons Paradox: Greater efficiency makes more uses worthwhile, expanding consumption and economic activity.
Roadmap:
- Deployment planned within OpenAI's infrastructure by year-end.
- Gen 2 is deep in development, and Gen 3 is taking shape.
More from Infra
- OpenAI Used Unreleased Model Astra to Design Its Jalapeño Inference Chip in 9 Months — jarrodwatts · 2026-08-26
- NVIDIA launches Jetson Orin Nano 2: 2x performance, 40% less power — nvidia · 2026-08-26
- Studies Reveal Data Centers Boost Local Jobs and Wages Significantly — justin_hart · 2026-08-26
- Opinion: Land Scarcity Drives Shift to Decentralized AI Compute — bittingthembits · 2026-08-26
- Jalapeno beats VR200 with optimized DeepSeek implementation, faster execution — itsclivetime · 2026-08-26
- CUDA code now runs on Apple Silicon with zero source code changes — petewoodbridge · 2026-08-26