Open-source Strata runs a 125B-param Qwen model on a 12GB consumer GPU
lxfater · x · 2026-10-03
The open-source project Strata (8k GitHub stars) runs Qwen3.8-Flash-Next, a 125-billion-parameter model, on an ordinary gaming PC — requiring only a 12GB+ NVIDIA/AMD GPU, with one-click install for Windows/Linux, a local OpenAI/Anthropic-compatible API, and optional image input.
It exploits MoE sparsity: hot experts live in VRAM while the rest are served from system RAM with CPU help, plus an SSD-backed lookup table for scheduling.
Benchmarks: on an RTX 5070 (12GB) + Ryzen 5 7600 + 64GB RAM, short-chat generation hits 94 token/s (Q20) and 53 token/s (IQ3S); another user (@ivanalogcom) with custom optimizations reached 2200 token/s prompt processing and 67 token/s generation on a single 5070 Ti. Quantization choice and context length significantly change speed and quality, so test on your own workload.
More from Infra
- Broadcom to lend Anthropic up to $42 billion to finance compute, IPO filings show — VraserX · 2026-10-04
- Workato's AI bill rose sevenfold as seat licences turned into usage meters — YvesMulkers · 2026-10-04
- Dev Slams Cloudflare D1 After 3 Months in Production: 3% of Reads Take 3-10s — TejasKumar_ · 2026-10-04
- Google's first orbital data center is in orbit: fridge-sized, 1 kW solar — lemire · 2026-10-03
- Japan's compute boom turns vending machine and toilet makers into hot AI stocks — PAstynome · 2026-10-03
- Chip design is a ~10^2,632,341 search problem — AI and agents are turning hardware into search — ai · 2026-10-03