OpenAI's VP of Hardware on Jalapeño: first custom AI chip taped out in nine months
ymatias · x · 2026-09-18
OpenAI VP of Hardware Richard Ho detailed the company's first custom AI accelerator Jalapeño on The Data Exchange podcast: AI tools helped compress parts of the development cycle to just nine months to tape-out. Key points:
- Why custom silicon: owning more of the stack to cut inference costs; a blank-slate architecture bringing memory closer to compute, breaking the throughput-vs-latency trade-off, targeting general-purpose LLMs and open-source models.
- Software stack: whether CUDA is still a moat, Triton, prefill/decode optimization, AI-optimized kernels, speculative decoding, plus InferenceX and AgentX benchmarks.
- Roadmap: Gen 1 to Gen 3 with HBM4 and new memory; he argues power, not chips, is becoming the data center bottleneck.
- AI-assisted engineering: increasingly capable AI tools could make expert engineering teams dramatically more productive.
Related event: OpenAI's first in-house AI chip Jalapeño taped out in 9 months(5 posts)→
More from Infra
- 421M-Parameter Laya Model Plays Flappy Bird on CPU via OpenVINO INT8 — simpleuserhere · 2026-09-23
- How to run Qwen3.8-27B with 160k context on a 16GB AMD card: full config — According_Study_162 · 2026-09-23
- MLX-Serve v26.9.5 lands with Qwen-Image 2.1 and 4-way MTP streams at up to 122 tok/s on M4 Max — TheMoonMidas · 2026-09-23
- DigitalOcean Managed Agents enters public preview: idle pausing, 75+ models, one bill — HeyAmit_ · 2026-09-23
- Swapping just the decision layer: Qwen + SGLang beats Jev by ~37% at same accuracy — VeryWellVersed · 2026-09-23
- Reka EdgeQ VLM Runs Natively on Snapdragon 8 Elite NPU With 0.73s First Token — RekaAILabs · 2026-09-23