Benchmarking Qwen3.8-Flash-Next + MTP on Strix Halo
betiz0 · reddit · 2026-08-29
A detailed technical report benchmarking Qwen3.8-Flash-Next combined with an MTP draft model on AMD Strix Halo hardware (Ryzen AI MAX+ 395 + Radeon 8060S).
The setup utilizes a specific fork of llama.cpp with Vulkan backend. Results show an average pre-fill speed (PP) of 138.61 tokens/s and a text generation speed (TG) of 26.67 tokens/s. The author notes that the output quality for architectural and design tasks appears noticeably superior to the 27B version, though it occasionally over-refines results.
More from Infra
- Qwen 350K Context Tested on M5 Max: Performance and Quality — Artistic_Okra7288 · 2026-08-30
- Azure Linux 4.0 Desktop Concept: PowerShell, Edge, and Copilot Pre-installed — unixterminal · 2026-08-30
- Jensen Huang: Built GPU tech first, found endless problems from graphics to molecular dynamics — r0ck3t23 · 2026-08-30
- How to build an LLM inference engine from scratch: 5-layer architecture — glenbeer · 2026-08-30
- Huaqin expects super node revenue to exceed 10B RMB in 2H 2026 — zephyr_z9 · 2026-08-30
- Nvidia is generating $1 billion a day, a business scale deemed absurd years ago — shauntrennery · 2026-08-30