OpenAI's new inference engine boosts throughput up to 4.1x on select models

rohanpaul_ai · x · 2026-08-25

Benchmarks show OpenAI's new inference engine 'Jalapeño' significantly boosts mixed-token throughput. It achieves 1,459 tokens/sec/user on GPT-OSS, 700 on DeepSeek R1, and 694 on Kimi K2.5. These figures represent increases of 2.7x, 4.1x, and 3.8x respectively compared to previous bests.

Original post →

More from Infra

Infra channel →