llama.cpp MoE host-memory cache PR underperforms on 8GB VRAM in user benchmarks
pmttyji · reddit · 2026-10-09
A user tested llama.cpp PR #29887 (GPU/host-memory cache for MoE experts) on a 4060 (8GB VRAM) + 32GB DDR5 with Qwen3.6-35B-A3B-IQ4XS. Plain -cmoe hit 23.4 t/s, but adding --moe-cache-mib made things worse: 15 t/s at 1536MiB, 13.4 t/s at 2048MiB, and 12 t/s at 8192MiB, with prompt eval also degrading. Unsure if it's misconfiguration or an implementation issue, the author asks the community for fixes and benchmarks from other low-VRAM users.
More from Infra
- Pixel 9A's TPU compiler so buggy that GPT trips security warnings; SDK under soft NDA — mgostIH · 2026-10-09
- Signal65: A model that fits in 128GB now matches Claude Opus 5 on agentic work, within 5% of a 2.4T flagship — ryanshrout · 2026-10-09
- Developer gets a full H100 node running at home — TheZachMueller · 2026-10-09
- DGX Spark prices skyrocket as resale markups soar — natesiggard · 2026-10-09
- OpenAI bots hit 160K fetches for nonexistent URLs in a week, sparking RL-run speculation — gaganghotra_ · 2026-10-09
- Modal's LLM Engine Advisor picks engine, model and config for your inference workload — charles_irl · 2026-10-09