Charles Frye (Modal) explains inference engines: schedulers, KV cache, CUDA graphs, speculative decoding
AI Engineer · youtube · 2026-10-06
Charles Frye of Modal gave an hour-long deep dive on inference engines at AI Engineer, opening with a museum placard generator whose first-token latency degrades under traffic spikes, and showing how to read queue pressure from a dashboard and relieve it with replicas.
Key content:
- Request lifecycle walkthrough: server IO → tokenization → scheduling → model execution → detokenization, using SGLang and vLLM as references
- The scheduler as hidden bottleneck: minimal computation but controls what reaches the GPU; chatbots, background agents, and document processors impose very different latency budgets, lengths, and prefix-reuse patterns
- Performance techniques: KV caching (capacity and layout drive throughput), CUDA graphs (cutting repeated CPU-side launch work), speculative decoding (a speculator proposes multiple tokens verified in one pass), kernel libraries and batch organization
- Production debugging: evaluate the real deployment, log token IDs to catch tokenizer issues, and collect metrics/traces to investigate regressions across replicas
He recommends starting with mini-sglang as a small reference engine for reading source code. Full timestamps and resources in the video description.
More from Infra
- Can a 128GB M5 Max Mac Studio Handle Concurrent Local LLM Agents? — Simple_Telephone_867 · 2026-10-06
- MiniMax discloses 70-80% inference margins, fueling debate on AI subscription subsidies — menhguin · 2026-10-06
- ~90% of frontier lab compute now goes to post-training and inference — IanAndrewsDC · 2026-10-06
- KLIF open-sources a Rust front-end that manages llama.cpp, vLLM and TTS servers in one window — Koksny · 2026-10-06
- Qwen 27B on 2× RX 7900 XT: 66.5 TPS single-stream, still short of claimed 100+ — EqualCryptographer67 · 2026-10-06
- Dev burns 842B tokens in September — $409k at API list price, pays just 3.4% via subscription — doodlestein · 2026-10-06