Test: Qwen 3.8 Flash Next on low-end hardware (6GB VRAM)

Agitated_Force_9199 · reddit · 2026-08-28

The author runs Qwen 3.8 Flash Next (1-bit quant) via llama.cpp on Ubuntu with 16GB RAM and 6GB VRAM, achieving a stable 6-7 tps. Impressed by the performance, they ask for advice on which quantization variant to try next for a balance of speed and quality.

Original post →

More from Infra

Infra channel →