Open-source Strata runs 125B MoE models on 12GB VRAM at up to 93 tokens/s
udmrzn · x · 2026-10-03
The open-source inference engine Strata runs 125B-class MoE models (Qwen3.8-Flash-Next) locally on consumer GPUs like an RTX 5070 12GB by splitting work across GPU, CPU, RAM, and SSD.
Key mechanisms:
- Doesn't load the whole model into VRAM: GPU holds Attention, Router, KV Cache, and hot experts; cold experts stay in RAM (CPU compute) with large lookup tables on SSD read on demand
- Adaptive expert caching lets 12–16GB VRAM handle huge MoE models
- Speculative decoding via the model's MTP layers, claimed 1.6–1.8x speedup
Benchmarks (RTX 5070 12GB + Ryzen 5 7600 + 64GB RAM, short context): 93 tokens/s at Q20, 79 at IQ2XS; at 128K context still 46–74 tokens/s. Recommended: 12GB+ VRAM, 64GB RAM, NVMe SSD, 70–120GB storage.
More from Infra
- Just 10% of people using AI heavily could strain inference infra as HBM costs skyrocket — XFreeze · 2026-10-03
- Critics Slam a "Neocloud": Wrapped RunPod Months Ago, No Colocated DCs — knowrohit07 · 2026-10-03
- Analyst Bullish on AAOI as 800G/1.6T Ramp; 3.2T Not Mainstream Yet — BenBajarin · 2026-10-03
- Google Dropped Its Tier 2 Spend Gate That Pushed a Dev to OpenRouter — vivekhaldar · 2026-10-03
- AMD: Software Tuning on MI355X Plus vLLM 0.30.1 Cuts MiniMax-M3 Token Cost by 31% — ryanshrout · 2026-10-03
- 4-GPU server topology: one Gen5 switch covers it, benchmark all-reduce before buying three — knowrohit07 · 2026-10-03