Qwen 3.8 27B on Mac hits 113 TPS via Dflash2 optimization

TheMoonMidas · x · 2026-08-25

Developer @TheDavidTai reported speed improvements for Qwen 3.8 27B on Mac by switching the previous MTP configuration to adaptive Dflash2 and porting mlx.fast changes. The model achieved 113 TPS on a 1024-token Python prompt. Benchmarks across 1k to 128k context lengths show an average 20% speed boost over the previous fastest config (10% from mlx.fast, 10% from Dflash2). Dflash2 generally outperforms MTP, except potentially in long thinking traces.

Original post →

More from Infra

Infra channel →