Baseten's Kiely: Speculative Decoding Is the Fastest-Moving Front in Inference

AI Engineer · youtube · 2026-09-20

Philip Kiely of Baseten reviews what changed in inference engineering since his book. The through-line: local inference is "get it working, then make it less dumb"; data center inference is "get it working, then make it less slow" — and the biggest shift is that data center optimizations increasingly come from dedicated training, blurring training and inference.

Original post →

More from Infra

Infra channel →