Full Talk Slides Released: How Inference Engines Actually Work, End to End
zainhas · x · 2026-09-12
zainhas released full slides for his talk "How Inference Engines Actually Work", covering the complete lifecycle of a request: engine architecture, KV + prefix caching, continuous batching, PagedAttention, chunked prefill, sampling/detokenization, and agentic loops from inside the engine. A systems-level primer for engineers on inference stacks like vLLM.
Related event: Full Slide Deck on How Inference Engines Work Released(2 posts)→
More from Infra
- ADSP Ep. 303: Mark Saroufim on Open vs Closed Models, AI GPU Kernels and Autoresearch — blelbach · 2026-09-12
- Repurpose an old low-VRAM GPU just for mmproj in llama.cpp — an order of magnitude faster — inthesearchof · 2026-09-12
- Ollama vs llama.cpp: same request counts 85k vs 164k tokens in OpenCode — RadianceTower · 2026-09-12
- A running list of 15+ VC funds actively investing in AI infrastructure — Sethwinterroth · 2026-09-12
- Zilliz CTO: Agent memory is a long-lived systems problem, not an index feature — J_Luan_ · 2026-09-12
- Draw Things update adds MiniMax H3 with LoRA/TeaCache and Krea 2 model imports — antirez · 2026-09-12