HF engineer releases open slide deck on local AI: quantization to speculative decoding
mervenoyann · x · 2026-10-07
Hugging Face's merve released a reusable (with attribution) slide deck on running AI locally with llama.cpp, covering prefill vs decode, MoE vs dense, VRAM vs unified memory, quantization, and speculative decoding — a systematic primer on local inference.
Related event: Hugging Face Engineer Open-Sources Local AI Deployment Slides(2 posts)→
More from Infra
- TablePlus launches a VM built for Apple Silicon and AI Agents with GPU acceleration and MCP — film_girl · 2026-10-07
- Realtime inference startup Reactor raises Series A led by Lightspeed, with Nvidia joining — buckymoore · 2026-10-07
- Terse, an open-source alternative to Cloudflare Durable Objects, joins YC F26 — mertdumenci · 2026-10-07
- Databricks launches Lakebase, a database rearchitected for AI-speed workloads — matei_zaharia · 2026-10-07
- Podcast: has the memory cycle broken, plus Nvidia & Broadcom's $40B financing playbook — BenBajarin · 2026-10-07
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07