Qwen3.8-Flash-Next Mac optimization: Linear sparse attention, SSD streaming, and custom Q4
memeka · reddit · 2026-08-30
Developer achieved extreme optimization for Qwen3.8-Flash-Next on M1 Max 64GB, enabling SSD streaming for tensors, engrams, and MTP. Key techniques include:
- Custom Q4 Quant: Spliced tensors from multiple Unsloth/AtomicChat quants for best performance.
- Metal Sparse Attention: Custom mechanism achieving near-linear degradation vs standard quadratic.
- Dynamic MTP: Automatically disables MTP when context size makes it counter-productive.
- Tradeoffs: Enabling MTP uses more RAM, dropping prefill from 180 tps to 170 tps, but boosting decode by 70% to 22 btps.
Author credits Claude for providing a week's worth of tokens and credits.
More from Infra
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01