Full talk slides released: how inference engines actually work, end to end
zainhas · x · 2026-09-12
zainhas gave a talk titled "how inference engines actually work" at the Presidio Bitcoin Open Source AI summit and released the full slide deck, covering the end-to-end lifetime of a request:
- the inference engine itself
- KV and prefix caching
- continuous batching
- Paged Attention
- chunked prefill
- sampling
- agentic loops from inside the engine
A solid resource for developers who want a systematic understanding of the LLM inference serving stack.
Related event: Full Slide Deck on How Inference Engines Work Released(3 posts)→
More from Infra
- Intel's silicon photonics couplers hit 1-1.5 dB IL, with visible epoxy delamination flaws — jwt0625 · 2026-09-12
- CPO paper criticized for vague DLW-to-PIC coupling description: 'such as TCB' — jwt0625 · 2026-09-12
- During AWS outage, one engineer kept enterprise services up with just 22 min downtime — generativist · 2026-09-12
- US hosts 43% of global datacenter power use; China just 13%, report finds — TMWNN · 2026-09-12
- Intel engineers show wafer-level chiplet testing for co-packaged optics paper — jwt0625 · 2026-09-12
- Intel engineers spotted testing chiplets and packages on wafer-level tester — jwt0625 · 2026-09-12