Strata runs 125B Qwen3.8-Flash-Next on a 16GB consumer GPU

evilsocket · x · 2026-10-01

evilsocket demos running Qwen3.8-Flash-Next (125B params, IQ1M quantization) on a single 16GB NVIDIA GPU with the open-source Strata inference engine. Highlights: one-click install on Windows/Linux needing a 12-24GB GPU plus 64GB RAM, 60-95 tokens/s output, local OpenAI/Anthropic-compatible APIs, optional image input, and 3.9k GitHub stars.

Related event: Running quantized Qwen3.8 coder on a 16GB GPU with Strata(2 posts)→

Original post →

More from Infra

Infra channel →