HF engineer releases free local AI slide deck: prefill vs decode, MoE, quantization, llama.cpp
mervenoyann · x · 2026-10-07
Hugging Face engineer Merve Noyan released her local AI slide deck, free to reuse with attribution.
It covers the core concepts of running models locally, all through the llama.cpp ecosystem:
- prefill vs decode inference stages
- MoE vs dense architectures
- VRAM vs unified memory
- quantization and speculative decoding
A systematic free primer for anyone getting into local deployment.
Related event: Hugging Face Engineer Open-Sources Local AI Deployment Slides(2 posts)→
More from Infra
- Databricks launches Lakebase, a database rearchitected for AI-speed workloads — matei_zaharia · 2026-10-07
- Podcast: has the memory cycle broken, plus Nvidia & Broadcom's $40B financing playbook — BenBajarin · 2026-10-07
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- HF engineer releases open slide deck on local AI: quantization to speculative decoding — mervenoyann · 2026-10-07
- 21M model + 6.4B SSD-resident lookup table matches a 114M dense model — fechyyy · 2026-10-07
- Ollama 0.35 adds local decision models from Cloudflare, Together and Bespoke — Technovangelist · 2026-10-07