Baseten’s inference masterclass maps the new stack for fast, reliable AI serving

Latent Space · rss · 2026-08-04

This Latent Space episode with Baseten’s Philip Kiely and Ali Taha argues that inference has become a standalone engineering discipline. It covers cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the economics of dedicated deployments versus shared APIs.

The conversation also goes beyond LLM serving into NVIDIA Dynamo, local inference, AI-specific chips, and the bottlenecks behind long-form video generation. A key theme is that optimization gains can still be huge: quantization can preserve benchmark quality while boosting throughput, and the team discusses how open models can be made fast and reliable enough for production from day zero.

Original post →

More from Infra

Infra channel →