Qwen 3.8 27B on Mac hits 113 TPS via Dflash2 optimization
TheMoonMidas · x · 2026-08-25
Developer @TheDavidTai reported speed improvements for Qwen 3.8 27B on Mac by switching the previous MTP configuration to adaptive Dflash2 and porting mlx.fast changes. The model achieved 113 TPS on a 1024-token Python prompt. Benchmarks across 1k to 128k context lengths show an average 20% speed boost over the previous fastest config (10% from mlx.fast, 10% from Dflash2). Dflash2 generally outperforms MTP, except potentially in long thinking traces.
More from Infra
- Strix Halo + dGPU real-world test: low-context benchmarks oversell the speedup — Hrethric · 2026-08-25
- Debunking the "zombie facts" haunting the data center debate — AndyMasley · 2026-08-25
- Fal releases post-trained H3 model co-optimized with custom inference stack — isidentical · 2026-08-25
- A 4060Ti 16GB running Qwen 27B at IQ3_XXS merged its first full feature branch — o0genesis0o · 2026-08-25
- Optimizing Qwen3-ASR Latency to 70ms to Beat Deepgram — Comprehensive_Quit67 · 2026-08-25
- Agent Memory System Refuses Hallucinations with Provable Deletion and Auditability — External-Fee-8920 · 2026-08-25