OpenAI reveals Jalapeno chip: B200-level throughput with novel speculate-decode split

appenz · x · 2026-08-26

OpenAI showcased its Jalapeno chip at HotChips. It offers throughput comparable to a B200 but with lower power consumption and higher efficiency at high token rates. The most notable feature is the separation of Speculate and Decode into distinct phases using different models, a novel architectural approach.

Related event: OpenAI Unveils Jalapeno Inference ASIC at Hot Chips(6 posts)→

Original post →

More from Infra

Infra channel →