OpenAI hardware VP details first custom chip Jalapeño and its nine-month tape-out
bigdata · x · 2026-09-20
The Data Exchange interviews Richard Ho, VP of Hardware at OpenAI, about Jalapeño, the company's first custom AI accelerator:
- Why custom: owning more of the stack to cut inference cost; unusual combo of high throughput and low latency, designed for general-purpose and open-source LLMs.
- Architecture: blank-slate design bringing memory and compute closer; prefill/decode, AI-optimized kernels, speculative decoding.
- AI-accelerated engineering: parts of the cycle compressed to a nine-month tape-out; roadmap reaches Gen 3 with HBM4 and new memory.
- Also covered: CUDA's moat, power becoming the datacenter bottleneck, RL/agents driving inference demand, InferenceX/AgentX benchmarks.
More from Infra
- Jevons paradox is classic low-end disruption you won't spot from a GPU-rich hyperlab — cramforce · 2026-09-20
- SpaceX's orbital AI data centers weigh up to 4,000 kg each, filing seeks 1M satellites — XFreeze · 2026-09-20
- Local Models for Personal Agents: GPT Luna Surprises a Coding-Agent Veteran — gized00 · 2026-09-20
- focus-llama: a llama.cpp fork implementing Declarative Attention for up to 0.71x decode time — Ok-Shower7286 · 2026-09-20
- VTrain, a Vulkan-based resident trainer, fixes memory leak and offloads more work to GPU — Savantskie1 · 2026-09-20
- vLLM ships day-0 support for Qwen-Image-2.1 with cross-step prefix KV cache — Alibaba_Qwen · 2026-09-20