OpenAI's new inference engine boosts throughput up to 4.1x on select models
rohanpaul_ai · x · 2026-08-25
Benchmarks show OpenAI's new inference engine 'Jalapeño' significantly boosts mixed-token throughput. It achieves 1,459 tokens/sec/user on GPT-OSS, 700 on DeepSeek R1, and 694 on Kimi K2.5. These figures represent increases of 2.7x, 4.1x, and 3.8x respectively compared to previous bests.
More from Infra
- Switching AI models frequently invalidates prompt cache, spiking costs — Daniel_Farinax · 2026-08-25
- Analyst: custom ASIC demand is very aggressive; upside for Qualcomm, AMD, Intel — BenBajarin · 2026-08-25
- MacStories: M6 and M5 Ultra offer huge potential for local AI on macOS — Dimillian · 2026-08-25
- Australia Faces Datacentre Rush as AI Firms Scramble to Dodge Upcoming Regulations — nordicinst · 2026-08-25
- OpenAI Claims New 'Jalapeño' Chip Outperforms Vera Rubin in Benchmarks — Wonderful_Buffalo_32 · 2026-08-25
- Jalapeño ASIC outperforms comparable TPUs, challenging existing giants — GavinSBaker · 2026-08-25