10 resources on what happens after training: KV-cache, quantization, serving
techNmak · x · 2026-09-11
A widely shared thread curates 10 resources explaining the often-ignored layer after model training: GPU memory, batching, KV-cache allocation, quantization, kernels, latency, and distributed serving. Commenters note that if you can't explain KV-cache and quantization, you're not done learning ML.
More from Infra
- Routing NVIDIA PAIR to llama.cpp on an AMD ROCm node (2×R9700): full notes — Don_Reuter · 2026-09-11
- Running Qwen3.8 locally on a 128GB laptop for agentic coding: thinking tokens, not tok/s, set the wall clock — deepu105 · 2026-09-11
- SpaceX plans to make scarce turbine parts as AI data centers outpace gas turbine supply — rohanpaul_ai · 2026-09-11
- AGI as task time horizon vs meetings — and why fabs should train their own models — jwt0625 · 2026-09-11
- Free client-side calculator compares LLM token economics across DeepSeek, Claude, o3-mini — nikola_mr64990 · 2026-09-11
- Full-precision DeepSeek 4.1 Flash hits 300+ TPS on 4 RTX Pros with custom vLLM fork — TheZachMueller · 2026-09-11