OpenAI's first custom inference chip Jalapeño claims 1.5-1.9x better perf-per-watt than Nvidia GB200/GB300
kimmonismus · x · 2026-08-25
OpenAI says its first custom inference chip Jalapeño already beats Nvidia GB200 and GB300 systems in its own InferenceX testing.
- Across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T: 1.5–1.9× more AI work per watt at peak throughput, 1.7–3.6× lower end-to-end latency
- Highly interactive workloads see 2.1–4.1× higher performance
- Rated at 700W but stayed at or below 550W during tests
- Deployment planned by end of 2026; Gen 2 deep in development, Gen 3 taking shape
The poster speculates this is why Tibo predicted 750 token/s becomes default in 1-2 years. Numbers come from OpenAI's own testing, pending independent verification.
More from Infra
- OpenAI's in-house inference chip reportedly rivals GB300, NVIDIA impact seen as limited — ivan_bezdomny · 2026-08-25
- Lambda seeks input on model cards: add NVFP4 weights and base models? — TheZachMueller · 2026-08-25
- Perplexity releases research on Portable Computer on Spark — AravSrinivas · 2026-08-25
- Arav Srinivas: On-device models critical for sensitive docs with SOTA OCR — AravSrinivas · 2026-08-25
- Arav Srinivas: Agentic inference must move to local hardware — AravSrinivas · 2026-08-25
- Data center boom drives a wave of gas power plant projects across the US — pstAsiatech · 2026-08-25