Qwen Flash Next Q4 hits 17.5 tks on Mac Mini M5 with carousel expert streaming and dual-SSD tricks

turtleninja99 · reddit · 2026-10-05

A Redditor details running Qwen Flash Next Q4 locally on a Mac Mini M5 (64GB), hitting 17.5 tks decode and 360 tks prompt processing. Beyond that, he shares optimizations few others have tried:

He argues there's plenty of headroom left — the GPU idles while experts stream in during decode, and an optimized Metal kernel could help. With Flash Next as a precursor to Qwen 4, he expects ngram tables and cheap hybrid attention caching to unlock more local-efficiency wins. He finds its coding quality surprisingly strong and leaves it running for large jobs, though long thinking time hurts real-world decode speed. Details and setup are open-sourced on GitHub (skeggsguy/Flash-next-ssd).

Original post →

More from coding & agent

coding & agent channel →