llama.cpp PR adds GPU cache for host-resident MoE experts, big speedup potential

jacek2023 · reddit · 2026-10-08

PR #29887 by am17an in ggml-org/llama.cpp adds a GPU cache for MoE experts kept in host memory, potentially delivering a significant speedup for MoE models that don't fully fit in VRAM and lowering the memory bar for local deployment.

Original post →

More from Infra

Infra channel →