Modular's LLM Inference Handbook covers prefill, KV cache, batching, quantization and speculative decoding
techNmak · x · 2026-09-04
A recommended resource: Modular's LLM Inference Handbook bridges "how LLMs work" and "how LLM systems serve requests," covering prefill, decode, KV cache, batching, scheduling, quantization, prefix caching and speculative decoding.
It also points to Maarten Grootendorst's Visual Guide to Mixture of Experts, which explains routers, experts, sparse activation, token routing and load balancing with dozens of visuals.
Related event: A Curated Thread of Visual and Interactive Resources for Learning AI(13 posts)→
More from Infra
- Kafka exactly-once delivery demystified: the 3-layer design and the catch engineers miss — arpit_bhayani · 2026-09-04
- AMD's Threadripper Halo Station packs 96 cores and 576GB of HBM3E — ccerrato147 · 2026-09-04
- Pennsylvania voters unite against data centres: 'People are going to get screwed' — 1vuio0pswjnm7 · 2026-09-04
- AeroJEPA fluid foundation model joins NVIDIA's PhysicsNeMo ecosystem — ricardovinuesa · 2026-09-04
- Building a €2-2.5k local AI rig for legal RAG and agentic coding: hardware picks debated — whatyathinkk · 2026-09-04
- Dual 3090 owners debate adding more cards: bigger local models vs parallel instances — Blues520 · 2026-09-04