LayerStoRm open-source engine runs 186GiB MoE on 96GB VRAM at 24.5 tok/s with 1M context

CharacterBumblebee99 · reddit · 2026-09-07

A Reddit user released LayerStoRm, an experimental MIT-licensed inference engine that streams MoE experts from pinned host RAM to GPU per token, letting you run models far larger than your VRAM.

Repo: github.com/kkontosis/LayerStoRm

Original post →

More from coding & agent

coding & agent channel →