Open-source Strata runs 125B MoE models on 12GB VRAM at up to 93 tokens/s

udmrzn · x · 2026-10-03

The open-source inference engine Strata runs 125B-class MoE models (Qwen3.8-Flash-Next) locally on consumer GPUs like an RTX 5070 12GB by splitting work across GPU, CPU, RAM, and SSD.

Key mechanisms:

Benchmarks (RTX 5070 12GB + Ryzen 5 7600 + 64GB RAM, short context): 93 tokens/s at Q20, 79 at IQ2XS; at 128K context still 46–74 tokens/s. Recommended: 12GB+ VRAM, 64GB RAM, NVMe SSD, 70–120GB storage.

Related event: Open-source engine Strata runs 125B Qwen3.8-Flash-Next on consumer gaming PCs(7 posts)→

Original post →

More from Infra

Infra channel →