Running Qwen3.8 Flash-Next locally on AMD 7900 XTX at 500k context, 105-160 tok/s

human_in_the_looop · reddit · 2026-10-12

A Redditor runs Qwen3.8 Flash-Next (176B total / 6B active MoE) locally on an AMD 7900 XTX with 500k context: 66.4GB weights split across 23.8GiB VRAM, 31.6GB expert weights and 28.8GB PLE table in system RAM, 105-160 tok/s decode, measured 506,849 tokens, quality 81-94% from 4k to 256k. Key gotcha: preservethinking: true caused an 'echo' loop (266 times in 13k turns) that self-reinforced; removing it fixed the issue.

Original post →

More from Infra

Infra channel →