Benchmarking Qwen3.8-Flash-Next + MTP on Strix Halo

betiz0 · reddit · 2026-08-29

A detailed technical report benchmarking Qwen3.8-Flash-Next combined with an MTP draft model on AMD Strix Halo hardware (Ryzen AI MAX+ 395 + Radeon 8060S).

The setup utilizes a specific fork of llama.cpp with Vulkan backend. Results show an average pre-fill speed (PP) of 138.61 tokens/s and a text generation speed (TG) of 26.67 tokens/s. The author notes that the output quality for architectural and design tasks appears noticeably superior to the 27B version, though it occasionally over-refines results.

Original post →

More from Infra

Infra channel →