Charles Frye (Modal) explains inference engines: schedulers, KV cache, CUDA graphs, speculative decoding

AI Engineer · youtube · 2026-10-06

Charles Frye of Modal gave an hour-long deep dive on inference engines at AI Engineer, opening with a museum placard generator whose first-token latency degrades under traffic spikes, and showing how to read queue pressure from a dashboard and relieve it with replicas.

Key content:

He recommends starting with mini-sglang as a small reference engine for reading source code. Full timestamps and resources in the video description.

Original post →

More from Infra

Infra channel →