Running Qwen3.8 Flash on 12GB VRAM at 15 tokens/s with 3bpw quantization
KnownAd4832 · reddit · 2026-09-17
A Redditor ran Qwen3.8-Flash-Next-GSQ-RCO-GGUF locally on a 12GB RTX 5070 SFF with 64GB DDR5, using a 3bpw quant (IQ3XXS, claimed to match BF16 on AIME25), achieving steady 15 tokens/s generation and 100-120 tokens/s prompt processing. The full 76GB GGUF needs only 47GB sharded into VRAM+RAM, the rest served from SSD.
Benchmarks: server startup 4.9s; cold launch to 1K prompt 31.7s; 1K generation 13.8 tok/s; 8K generation 14.6 tok/s; 20K prompt processing 104.7 tok/s (3.11 min wait); 20K generation 14.0 tok/s; 90–100K generation 11.3 tok/s; verified up to 128K context. Output quality reportedly matches Unsloth's Q6–Q8 level; built with FreeToken CLI + llama; a 2.40bpw version (targeting NVFP4/Q4 parity) is next.
More from Infra
- AWS shows NVRx fault-tolerant FSDP training on EKS: sync checkpointing ate up to 40% of wall time — AWS ML Blog · 2026-09-17
- Qwen 3.8 27B NVFP4 benchmarked on 2xV100 with ~400k context — jjusko20 · 2026-09-17
- NVIDIA launches CUDA Rust: two paths to write GPU kernels natively in Rust — blelbach · 2026-09-17
- AI agent ports and optimizes CUDA attention kernel to CuTeDSL in an afternoon, 1.2x faster — knowrohit07 · 2026-09-17
- llama.cpp PR Enables CUDA Graphs for MTP Draft, Delivering Another Inference Speedup — jacek2023 · 2026-09-17
- Open-Source Maestro: One-Click Local AI Creative Studio for Music, Video and More — cocktailpeanut · 2026-09-17