LayerStoRm: Run Frontier-Scale MoE LLMs on Consumer GPUs via PCIe Streaming
CharacterBumblebee99 · reddit · 2026-08-25
LayerStoRm is an open-source MoE LLM serving engine for limited VRAM multi-GPU setups, utilizing RAM and PCIe transfers. It optimizes per-layer transfer schedules, fetching only the experts routed by each token. Currently targeting RTX 5090/5080 systems, it achieves 10 gen tok/sec with GLM 5.2. Features include KV offloading, PagedAttention, and OpenAI-compatible API. It is an experimental, research-grade release under the MIT license.
More from Infra
- AI Supply Chain Faces Bullwhip Effect, HDD Prices Surge — AccBalanced · 2026-08-25
- From Notebook to Production: A 15-Day MLOps Learning Roadmap — _jaydeepkarale · 2026-08-25
- Debunking Data Center Myths: Water, Power, Taxes, and Land Use — AndyMasley · 2026-08-25
- Strix Halo + dGPU real-world test: low-context benchmarks oversell the speedup — Hrethric · 2026-08-25
- Fal releases post-trained H3 model co-optimized with custom inference stack — isidentical · 2026-08-25
- A 4060Ti 16GB running Qwen 27B at IQ3_XXS merged its first full feature branch — o0genesis0o · 2026-08-25