Qwen 3.5 9B + DFlash hits 75 tok/s on R9700: Full Setup Guide

karmakaze1 · reddit · 2026-08-20

A Reddit user shared a detailed guide to achieving 75 tokens/sec inference speed with the Qwen 3.5 9B model on an R9 7900 GPU. The setup combines a fine-tuned model (DavidAU/Qwen3.5-9B-The-Defiant-Fable...) with DFlash draft model acceleration, utilizing Q80 quantization and 64K context.

Key steps include:

This approach leverages a draft model to significantly boost generation throughput.

Original post →

More from coding & agent

coding & agent channel →