Offloading MoE models to RAM causes slow prefill speeds

former_farmer · reddit · 2026-08-24

Community discussions highlight that while offloading MoE models to system RAM allows decent decode speeds on GPUs with low VRAM, the prefill phase becomes extremely slow, significantly degrading the user experience. Users are confirming if this performance bottleneck is consistent.

Original post →

More from Infra

Infra channel →