HF engineer releases open slide deck on local AI: quantization to speculative decoding

mervenoyann · x · 2026-10-07

Hugging Face's merve released a reusable (with attribution) slide deck on running AI locally with llama.cpp, covering prefill vs decode, MoE vs dense, VRAM vs unified memory, quantization, and speculative decoding — a systematic primer on local inference.

Related event: Hugging Face Engineer Open-Sources Local AI Deployment Slides(2 posts)→

Original post →

More from Infra

Infra channel →