Baseten says inference optimizations can still make open models up to 10× faster
Latent Space · youtube · 2026-08-04
Inference is becoming the new frontier
Baseten’s Philip Kiely and Ali Taha walk through what actually happens after an open model is released and how a model becomes a production API.
- The episode focuses on cache-aware routing, disaggregated prefill/decode, quantization, speculative decoding, KV-cache movement, model parallelism, and GPU kernels.
- They argue inference optimizations can still yield 20%–200% gains, and in some setups make models up to 10× faster.
- The discussion covers serving day-one open models, mixing components across architectures, why identical weights can behave differently across clusters, and how models can even help optimize the infrastructure that serves them.
- It also touches NVIDIA Dynamo, Rubin, local inference, AI-specific chips, and the compute barriers to coherent long-form video generation.
Related event: Baseten Team Shares Insights on AI Inference Optimization(3 posts)→
More from coding & agent
- A prompt pack for making Claude Code write clearer technical reports — wzenus · 2026-08-04
- Claude 3.5 Sonnet Aces Adversarial Data Science Test, Catches Data Leakage Autonomously — hugobowne · 2026-08-04
- GitHub list tracks OSINT MCP servers for Claude, Cursor and Windsurf — tom_doerr · 2026-08-04
- Vibe coding gives old tools like MediaPipe a second life — bilawalsidhu · 2026-08-04
- Google details the pipeline behind its open-source Agent Skills — _jaydeepkarale · 2026-08-04
- Steve Yegge says Opus 4.7 kept tweaking Gas Town instead of finishing work — Simon Willison · 2026-08-04