Open-source Strata runs a 125B-param Qwen model on a 12GB consumer GPU

lxfater · x · 2026-10-03

The open-source project Strata (8k GitHub stars) runs Qwen3.8-Flash-Next, a 125-billion-parameter model, on an ordinary gaming PC — requiring only a 12GB+ NVIDIA/AMD GPU, with one-click install for Windows/Linux, a local OpenAI/Anthropic-compatible API, and optional image input.

It exploits MoE sparsity: hot experts live in VRAM while the rest are served from system RAM with CPU help, plus an SSD-backed lookup table for scheduling.

Benchmarks: on an RTX 5070 (12GB) + Ryzen 5 7600 + 64GB RAM, short-chat generation hits 94 token/s (Q20) and 53 token/s (IQ3S); another user (@ivanalogcom) with custom optimizations reached 2200 token/s prompt processing and 67 token/s generation on a single 5070 Ti. Quantization choice and context length significantly change speed and quality, so test on your own workload.

Original post →

More from Infra

Infra channel →