Strata runs 125B Qwen3.8-Flash-Next on a 16GB consumer GPU
evilsocket · x · 2026-10-01
evilsocket demos running Qwen3.8-Flash-Next (125B params, IQ1M quantization) on a single 16GB NVIDIA GPU with the open-source Strata inference engine. Highlights: one-click install on Windows/Linux needing a 12-24GB GPU plus 64GB RAM, 60-95 tokens/s output, local OpenAI/Anthropic-compatible APIs, optional image input, and 3.9k GitHub stars.
Related event: Running quantized Qwen3.8 coder on a 16GB GPU with Strata(2 posts)→
More from Infra
- DeepSeek V4.1 Flash spotted running locally on a 192GB Framework Desktop — antirez · 2026-10-01
- China's CXMT to nearly match Micron's DRAM capacity by end of 2026 — Terminator857 · 2026-10-01
- FreeToken vs llama.cpp on RTX 3090: 7x faster TTFT only when the MoE won't fit in VRAM — SignatureMoney6648 · 2026-10-01
- Building a million-page OCR pipeline with a 500GB RAM used server plus LLM extraction — oilmutt · 2026-10-01
- Volantis raises $88M to break the memory wall with photonics, eyeing 10,000 tok/s inference — dunkhippo33 · 2026-10-01
- Corbenic AI launches Galahad beta: disk-persistent KV cache reuse across requests and restarts — MindPsychological140 · 2026-10-01