LayerStoRm: Run Frontier-Scale MoE LLMs on Consumer GPUs via PCIe Streaming

CharacterBumblebee99 · reddit · 2026-08-25

LayerStoRm is an open-source MoE LLM serving engine for limited VRAM multi-GPU setups, utilizing RAM and PCIe transfers. It optimizes per-layer transfer schedules, fetching only the experts routed by each token. Currently targeting RTX 5090/5080 systems, it achieves 10 gen tok/sec with GLM 5.2. Features include KV offloading, PagedAttention, and OpenAI-compatible API. It is an experimental, research-grade release under the MIT license.

Original post →

More from Infra

Infra channel →