llama.cpp MoE host-memory cache PR underperforms on 8GB VRAM in user benchmarks

pmttyji · reddit · 2026-10-09

A user tested llama.cpp PR #29887 (GPU/host-memory cache for MoE experts) on a 4060 (8GB VRAM) + 32GB DDR5 with Qwen3.6-35B-A3B-IQ4XS. Plain -cmoe hit 23.4 t/s, but adding --moe-cache-mib made things worse: 15 t/s at 1536MiB, 13.4 t/s at 2048MiB, and 12 t/s at 8192MiB, with prompt eval also degrading. Unsure if it's misconfiguration or an implementation issue, the author asks the community for fixes and benchmarks from other low-VRAM users.

Original post →

More from Infra

Infra channel →