Test: Qwen 3.8 Flash Next on low-end hardware (6GB VRAM)
Agitated_Force_9199 · reddit · 2026-08-28
The author runs Qwen 3.8 Flash Next (1-bit quant) via llama.cpp on Ubuntu with 16GB RAM and 6GB VRAM, achieving a stable 6-7 tps. Impressed by the performance, they ask for advice on which quantization variant to try next for a balance of speed and quality.
More from Infra
- Nvidia Q2 revenue hits $30B, future commitments skyrocket to $366B — GaryMarcus · 2026-08-28
- Acorn: AI Chief of Staff running on your own server — MatthewChang · 2026-08-28
- Trade unions push back against data center opposition — MatthewBerman · 2026-08-28
- UK Labour rejects Green party call to halt AI datacentre construction — nordicinst · 2026-08-28
- a16z Raises $1.1B Machine Age Fund for AI Physical Infrastructure — appenz · 2026-08-28
- llama.cpp mmap fits Qwen3.8-Flash-Next in 16G+64G RAM at 26t/s — q8019222 · 2026-08-28