Running Qwen3.8 Flash on 12GB VRAM at 15 tokens/s with 3bpw quantization

KnownAd4832 · reddit · 2026-09-17

A Redditor ran Qwen3.8-Flash-Next-GSQ-RCO-GGUF locally on a 12GB RTX 5070 SFF with 64GB DDR5, using a 3bpw quant (IQ3XXS, claimed to match BF16 on AIME25), achieving steady 15 tokens/s generation and 100-120 tokens/s prompt processing. The full 76GB GGUF needs only 47GB sharded into VRAM+RAM, the rest served from SSD.

Benchmarks: server startup 4.9s; cold launch to 1K prompt 31.7s; 1K generation 13.8 tok/s; 8K generation 14.6 tok/s; 20K prompt processing 104.7 tok/s (3.11 min wait); 20K generation 14.0 tok/s; 90–100K generation 11.3 tok/s; verified up to 128K context. Output quality reportedly matches Unsloth's Q6–Q8 level; built with FreeToken CLI + llama; a 2.40bpw version (targeting NVFP4/Q4 parity) is next.

Original post →

More from Infra

Infra channel →