Qwen Flash Next Q4 hits 17.5 tks on Mac Mini M5 with carousel expert streaming and dual-SSD tricks
turtleninja99 · reddit · 2026-10-05
A Redditor details running Qwen Flash Next Q4 locally on a Mac Mini M5 (64GB), hitting 17.5 tks decode and 360 tks prompt processing. Beyond that, he shares optimizations few others have tried:
- Carousel buffer for expert streaming during prompt processing: +30% pp throughput
- Second external SSD for parallel reads: +15% on both pp and decode
- Job eviction setup: a chat app can kick a long coding run to the background and resume it after chatting, useful for saving local cache memory
He argues there's plenty of headroom left — the GPU idles while experts stream in during decode, and an optimized Metal kernel could help. With Flash Next as a precursor to Qwen 4, he expects ngram tables and cheap hybrid attention caching to unlock more local-efficiency wins. He finds its coding quality surprisingly strong and leaves it running for large jobs, though long thinking time hurts real-world decode speed. Details and setup are open-sourced on GitHub (skeggsguy/Flash-next-ssd).
More from coding & agent
- Human-AI collaboration pushes 11-square packing lower bound to 3.875, closing 95.89% of the gap — ctjlewis · 2026-10-05
- Follow-up on the same 11-square packing breakthrough: 300-line verifiable proof via 1,121 weighted dots — ctjlewis · 2026-10-05
- Mobilerun scores perfect 116/116 on Google's AndroidWorld, beating Artemis — SimplyAnnisa · 2026-10-05
- Gemini Robotics-ER 2 plans dual Franka FR3 manipulation, open-sourced with Isaac Sim — Stefania_druga · 2026-10-05
- Dimillian: as models get smarter, intent-based single-thread prompting beats agent orchestration — Dimillian · 2026-10-05
- Dev builds an AI answering machine for his kids with Astra, 3D printing and an ESP32 — Stefania_druga · 2026-10-05