Qwen3.8 IQ1_S on RTX 5070: Achieving 22t/s with 12GB VRAM
jacek2023 · reddit · 2026-08-29
User reports running Qwen3.8-Flash-Next with extreme IQ1S quantization on a single RTX 5070 (12GB VRAM).
Setup & Command
- Model: Qwen3.8-Flash-Next-UD-IQ1S (GGUF)
- Hardware: Single RTX 5070
- Args: --parallel 1, -c 10000
Performance Metrics
- Generation speed: 21-22 tokens/s
- Prompt processing: 25 tokens/s
- VRAM: Fits within 12GB
The user recommends starting with small quantizations for low-VRAM setups to verify compatibility.
More from Infra
- TPU Origin Story: Google Speech Success Disaster Created Hardware Bottleneck — demian_ai · 2026-08-29
- Tips for local video generation on 16GB VRAM — mxjxn · 2026-08-29
- Tencent compresses Hy4-preview from 1.5TB to 200GB GGUF — RedditUsr2 · 2026-08-29
- Tips for local video generation on 16GB VRAM — PurzBeats · 2026-08-29
- Exo Labs claims 4.8 TB/s memory bandwidth from clustered Mac Studios, scaling linearly — anonmt57 · 2026-08-29
- Solo researcher boosts Qwen 3.8 on Mac: +54% decode at 147k context — julianharris · 2026-08-29