Users compare CPU/RAM offload setups for large models and mixed MoE workloads

Jorlen · reddit · 2026-07-29

A Reddit user asks how people run large models across CPU/RAM and GPU layers, especially with MoE models that don’t seem to behave well when many layers are shared from VRAM to system memory.

They’re collecting real-world specs and performance data, including:

The goal is to understand whether upgrading RAM is worthwhile and whether the observed performance issues are due to software configuration or specific hardware setups.

Original post →

More from Infra

Infra channel →