Baseten’s inference masterclass maps the new stack for fast, reliable AI serving
Latent Space · rss · 2026-08-04
This Latent Space episode with Baseten’s Philip Kiely and Ali Taha argues that inference has become a standalone engineering discipline. It covers cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the economics of dedicated deployments versus shared APIs.
The conversation also goes beyond LLM serving into NVIDIA Dynamo, local inference, AI-specific chips, and the bottlenecks behind long-form video generation. A key theme is that optimization gains can still be huge: quantization can preserve benchmark quality while boosting throughput, and the team discusses how open models can be made fast and reliable enough for production from day zero.
More from Infra
- K3 looks stronger than expected, and the argument is to own your inference stack — hsu_byron · 2026-08-04
- Photon 2.0 compiles Moondream, Qwen 3.5 and Gemma 4 into megakernels — sloppenheimer · 2026-08-04
- Minimax H3 runs on an RTX 3060 12GB, but a 10-second clip takes 30 minutes — irmemon225 · 2026-08-04
- Minimax H3 generation times on a 5070 Ti compare ComfyUI, Sage Attention, and Easy Cache — TheRedHairedHero · 2026-08-04
- Open-source H3Zero wraps MiniMax H3 in a Modal-hosted UI with API support — LegacyV1 · 2026-08-04
- MiniMax H3 already handles 15-second clips, but one RTX 5090 run took over 6 minutes — Careless-Constant-33 · 2026-08-04