Reddit Petitions llama.cpp for Hot Expert Reload to Speed Up Local MoE Inference

perelmanych · reddit · 2026-09-12

A Reddit user is asking llama.cpp maintainers to implement Hot Expert Reload on GPU, arguing it would meaningfully boost decode speed for MoE models with moderate active parameters (Qwen3.8-Flash-Next, DeepSeek V4/V4.1 Flash, GLM 5.3 Flash) even on a single 3090 — and approach full VRAM offload speeds with two. That would make these near-SOTA models genuinely usable locally.

Original post →

More from Infra

Infra channel →