Challenge boosts Qwen 3.8 speed by 152% on Mac
gajesh · x · 2026-08-16
The Qwen MLX Challenge hosted by MLX.fast has achieved a breakthrough, demonstrating the potential of Apple Silicon for running dense models.
Optimization Results
- The current top score improved Qwen 3.8 27B inference speed on Mac by 152.5% over the baseline.
- Decoding speed reached 53.7 tokens/s, 2.5x faster than native multi-token prediction (MTP).
Technical Details
- Speculative decoding is key for dense models; solvers achieved gains by optimizing draft head weights.
- The scoring mechanism was updated to the median across eight hidden lengths to prevent overfitting.
- This result challenges the widespread belief that Macs are slow for dense models.
Related event: MLX Community Challenge Boosts Qwen 3.8 Speed on Apple Silicon by Over 150%(2 posts)→
More from Infra
- Qwen 27B hits nearly 100 tokens/s on a local RTX 4090 via Ollama and Pinokio — cocktailpeanut · 2026-08-16
- Nvidia reportedly investing $3B in SB Energy to back OpenAI data centers — rohanpaul_ai · 2026-08-16
- Seeed Unveils reComputer RK3576 Edge AI Module — ___Mufasaa · 2026-08-16
- Agent Capacity Planning Guide: Avoiding production surprises — blaizedsouza · 2026-08-16
- How GPU Architecture and Memory Bandwidth Dictate LLM Inference Speed — blaizedsouza · 2026-08-16
- Apple MLX Ecosystem Fragmented, Needs Leadership — andrejusb · 2026-08-16