Qwen 177B at 11-15 tok/s on a Single RTX 5070 12GB: Expert Streaming Deep Dive
ayobluestarr · reddit · 2026-10-03
A Redditor ran Qwen3.8-Flash-Next 177B (UD-IQ3XXS) locally on an RTX 5070 12GB + 32GB DDR4 + Ryzen 5 5600GT with a custom llama.cpp expert-streaming setup, hitting 11.5 tok/s on benchmarks (up from 7), 14-15 tok/s in chat, and generating a working single-file Snake game at 10.15 tok/s over 4,892 tokens.
Key optimizations:
- Fixed Windows I/O queue-depth issues; one file handle per worker
- Built a page-locked hot-expert tier so the GPU pulls hot expert weights faster
- Quality-gated outputs against a control model; benchmarks used a heat file from a separate prompt set
The GitHub repo documents failed attempts, benchmark scripts, and methodology; the author is soliciting streaming suggestions.
More from Infra
- Inference now bigger than training, and it's reshaping data center buildouts — AccBalanced · 2026-10-03
- Terafab reportedly targets 1 TW of compute per year, ~50x current global AI output — XFreeze · 2026-10-03
- Local LLM inference speeds jump ~10x in a week: single consumer GPU now hits 2200 prefill — oran_ge · 2026-10-03
- Micron claims NVHBM will boost margins, but critics ask who captures the custom HBM value — AccBalanced · 2026-10-03
- Samsung claims first working die on sub-10nm 10a DRAM, mass production eyed for 2028 — zephyr_z9 · 2026-10-03
- Fake Intel N150 mini PC scam exposed: seller hardcoded 'New_N150' into BIOS string — Ok-Shower7286 · 2026-10-03