Open-source engine Strata runs 125B Qwen3.8-Flash-Next on consumer gaming PCs
Strata, an open-source inference engine built by a solo developer, lets the 125-billion-parameter Qwen3.8-Flash-Next run locally on an ordinary gaming PC. It racked up 5000+ GitHub stars within 8 days of launch, making it the focal project in local inference this week. Multiple community members have already verified it works in practice, and the barrier to entry is far lower than expected.
Confirmed
- Strata supports one-click installation on Windows/Linux — setup is just double-clicking a single file — and can plug into Claude Code (@alexverem)
- It runs the 125B-parameter Qwen3.8-Flash-Next on consumer hardware and serves a local OpenAI/Anthropic-compatible API (@alexverem)
- Security researcher evilsocket demonstrated running the IQ1M extreme-quantized qwen3.8-flash-next-coder on a single 16GB NVIDIA GPU, saying the hardware threshold is roughly a 12–24GB VRAM card (@evilsocket)
- Japanese user CurieuxExplorer tested it: the 87GB Qwen3.8 Flash Next ran at 120–150 t/s on a single RTX 5090 with Prefill at 4K t/s, hooked up to LibreChat (@CurieuxExplorer)
Why it matters
- Local deployment of large models usually depends on high-end servers or multi-GPU clusters; Strata's extreme quantization and inference optimizations push the threshold down to a consumer card with 16GB of VRAM, which has real value for privacy-sensitive scenarios and offline development workflows
- The project provides an OpenAI/Anthropic-compatible API, meaning existing AI toolchains like Claude Code and LibreChat can switch directly to local models at very low cost
- Independent tests from a security researcher to ordinary users confirm usability and throughput, suggesting the performance is not just marketing numbers under ideal conditions
2026-10-01 ~ 2026-10-03 · 7 related posts
Primary sources
- [source] 87GB Qwen3.8 Flash Next runs at 120-150 t/s on a single RTX 5090 — CurieuxExplorer · 2026-10-01
- [source] Running quantized Qwen3.8 coder on a 16GB GPU with Strata — evilsocket · 2026-10-01
- Strata runs 125B Qwen3.8-Flash-Next on a 16GB consumer GPU — evilsocket · 2026-10-01
- [source] Strata open-source engine runs 125B Qwen3.8-Flash-Next on gaming PCs, 5k stars in 8 days — alex_verem · 2026-10-02
- Strata repo: 125B Qwen3.8-Flash-Next on consumer hardware with local OpenAI/Anthropic APIs — alex_verem · 2026-10-02
- Strata runs a 125B model at 70tps on old DDR3 PCs with a cheap GPU — udmrzn · 2026-10-03
- Open-source Strata runs 125B MoE models on 12GB VRAM at up to 93 tokens/s — udmrzn · 2026-10-03