Developers Request 3-Tier MoE Offloading (Disk/CPU/GPU) for Local LLMs

storm1er · reddit · 2026-08-06

A developer submitted a feature request asking local LLM runners (like llama.cpp) to support arguments like --disk-moe.

The goal is to enable 3-tier MoE (Mixture of Experts) offloading across GPU, CPU, and Disk. This would significantly improve the feasibility of running large MoE models on machines with limited VRAM and RAM, resonating with local AI enthusiasts.

Original post →

More from Infra

Infra channel →