OpenAI's First Custom AI Chip Jalapeño Taped Out in Nine Months, Targets Low-Latency Inference
bigdata · x · 2026-09-18
Richard Ho, VP of Hardware at OpenAI, details Jalapeño, the company's first custom AI accelerator, in a podcast interview.
- Nine-month tape-out: AI tools helped compress parts of the development cycle, shrinking the design-to-tape-out timeline to nine months.
- Core motivation: Owning more of the hardware stack to lower inference costs and improve performance, rather than chasing peak compute.
- Architecture: A blank-slate design that brings memory and compute closer together, breaking the traditional throughput-vs-latency trade-off and targeting general-purpose LLMs and open-source models.
- Bottleneck: Power, not chips, is becoming the data center constraint.
- Software stack: CUDA, Triton, speculative decoding, kernel optimization, plus cross-model benchmarks like InferenceX and AgentX.
- Roadmap: From gen one to gen three, involving HBM4 and new memory; RL and agents are driving inference demand.
- CUDA moat: Discussion of whether CUDA remains a moat and the optionality created by AI and new hardware.
Related event: OpenAI's first in-house AI chip Jalapeño taped out in 9 months(5 posts)→
More from Infra
- 421M-Parameter Laya Model Plays Flappy Bird on CPU via OpenVINO INT8 — simpleuserhere · 2026-09-23
- How to run Qwen3.8-27B with 160k context on a 16GB AMD card: full config — According_Study_162 · 2026-09-23
- MLX-Serve v26.9.5 lands with Qwen-Image 2.1 and 4-way MTP streams at up to 122 tok/s on M4 Max — TheMoonMidas · 2026-09-23
- DigitalOcean Managed Agents enters public preview: idle pausing, 75+ models, one bill — HeyAmit_ · 2026-09-23
- Swapping just the decision layer: Qwen + SGLang beats Jev by ~37% at same accuracy — VeryWellVersed · 2026-09-23
- Reka EdgeQ VLM Runs Natively on Snapdragon 8 Elite NPU With 0.73s First Token — RekaAILabs · 2026-09-23