llama.cpp PR adds GPU cache for host-resident MoE experts, big speedup potential
jacek2023 · reddit · 2026-10-08
PR #29887 by am17an in ggml-org/llama.cpp adds a GPU cache for MoE experts kept in host memory, potentially delivering a significant speedup for MoE models that don't fully fit in VRAM and lowering the memory bar for local deployment.
More from Infra
- GitHub Copilot will offload tasks to local models like MAI-Code-1.1 Flash via HydraFusion — mariorod1 · 2026-10-08
- Microsoft launches Surface Spark with Nvidia: $6k for 128GB, odd form factor — casper_hansen_ · 2026-10-08
- Google launches first test satellite carrying 4 TPUs for its space-based ML infrastructure moonshot — CurieuxExplorer · 2026-10-08
- GPU rental platform Lium hits all-time-high utilization, courts idle GPU owners with 95%+ revenue share — const_reborn · 2026-10-08
- Marvell details Google chip deal worth up to $120B as inference accelerators split from TPUs — demian_ai · 2026-10-08
- llama.cpp Takes the Stage at Microsoft's Windows Event, Creator Celebrates Local AI Milestone — ggerganov · 2026-10-08