Running Qwen3.8 Flash-Next locally on AMD 7900 XTX at 500k context, 105-160 tok/s
human_in_the_looop · reddit · 2026-10-12
A Redditor runs Qwen3.8 Flash-Next (176B total / 6B active MoE) locally on an AMD 7900 XTX with 500k context: 66.4GB weights split across 23.8GiB VRAM, 31.6GB expert weights and 28.8GB PLE table in system RAM, 105-160 tok/s decode, measured 506,849 tokens, quality 81-94% from 4k to 256k. Key gotcha: preservethinking: true caused an 'echo' loop (266 times in 13k turns) that self-reinforced; removing it fixed the issue.
More from Infra
- $2 ESP32 board runs Pi-hole-style DNS blocker with 140K domains in 0.7MB — LinusEkenstam · 2026-10-12
- Engineer joins NVIDIA's Groq LPU compilers team, working on multi-chip partitioning — blelbach · 2026-10-12
- Linus Ekenstam wants nothing less than a 100B-param model running on your phone — LinusEkenstam · 2026-10-12
- AI buildout to cost $10.3 trillion to finance through 2032, topping all prior US investment booms — KyeGomezB · 2026-10-12
- Fireworks: open models plus fine-tuning match closed ones — Cursor gets 13x faster inference — AI Engineer · 2026-10-12
- Running a 456GB model on 192GB VRAM: offloaded inference hits 60-125 tok/s with 1M context — HankYeomans · 2026-10-12