Full Talk Slides Released: How Inference Engines Actually Work, End to End

zainhas · x · 2026-09-12

zainhas released full slides for his talk "How Inference Engines Actually Work", covering the complete lifecycle of a request: engine architecture, KV + prefix caching, continuous batching, PagedAttention, chunked prefill, sampling/detokenization, and agentic loops from inside the engine. A systems-level primer for engineers on inference stacks like vLLM.

Related event: Full Slide Deck on How Inference Engines Work Released(2 posts)→

Original post →

More from Infra

Infra channel →