OpenAI's Custom Inference Chip Jalapeño Delivers 1.9x More Work Per Watt Than Nvidia GB300
ivan_bezdomny · x · 2026-08-26
OpenAI's first custom inference chip 'Jalapeño' reportedly delivers 1.5-1.9x more work per watt and 1.7-3.6x lower latency than Nvidia's GB300 system. Built on TSMC N3P with 15.4TB/s bandwidth, the 700W chip was designed using AI assistance and is scheduled for deployment by late 2026.
Related event: OpenAI Unveils First Custom Inference Chip Jalapeño with Benchmark Results(42 posts)→
More from Infra
- InferenceX 3.x adds AgentX workloads, demanding long-context tests from next-gen accelerators — AccBalanced · 2026-08-26
- Firewall tuning becomes essential to combat bot scraping costs — mobileraj · 2026-08-26
- AI workflow costs to drop 95% in 6 months with OS models and new infrastructure — JosephJacks_ · 2026-08-26
- Sam Altman Hints at OpenAI's Custom Chip: 'It is Fast' — LinusEkenstam · 2026-08-26
- How to run local LLMs on low-end hardware (i7-8700 + RTX 2060)? — pet3121 · 2026-08-26
- Texas signs tax agreement, advancing SpaceX's TeraFab chip plant to agreement phase — elonmusk · 2026-08-26