Reddit Petitions llama.cpp for Hot Expert Reload to Speed Up Local MoE Inference
perelmanych · reddit · 2026-09-12
A Reddit user is asking llama.cpp maintainers to implement Hot Expert Reload on GPU, arguing it would meaningfully boost decode speed for MoE models with moderate active parameters (Qwen3.8-Flash-Next, DeepSeek V4/V4.1 Flash, GLM 5.3 Flash) even on a single 3090 — and approach full VRAM offload speeds with two. That would make these near-SOTA models genuinely usable locally.
More from Infra
- Viral Post: '"We Have the Compute" Was a Lie' — julian88888888 · 2026-09-12
- Open-source Deplo self-hosted deploy platform: no Docker, SSH or YAML, AI agents included — deplocloud · 2026-09-12
- Open-source models account for just ~3.2% of global AI training spend, analysis finds — 0xBekket · 2026-09-12
- OpenAI CFO: 'People are chasing GPUs — agentic AI is going to activate CPUs' — rohanpaul_ai · 2026-09-12
- 512GB DDR4 + 2x RTX 3090: What Local Models Should You Run on This Setup? — ArtifartX · 2026-09-12
- Meta Gives Every User a Free 2-vCPU VM With 100GB Storage, Compute Moat Even OpenAI Can't Match — signulll · 2026-09-12