mlx.fast fixes speed-display bug: MLX kernels hit 80.6 tps, nearing 100% speedup milestone
HankYeomans · x · 2026-09-17
TheDavidTai, author of the mlx.fast inference-optimization project, fixed two bugs:
- Kernels are much faster than shown: an infrastructure rewrite caused decode speed to be divided by total runtime, understating performance. Corrected measurements show 80.6 tps on MLX and 43.5 tps on CUDA. He credits skeptical community members for surfacing the issue.
- Yukon points were mis-awarded: points weren't granted properly; fixed going forward.
The leaderboard's current record stands at a 97.3% composite speedup over the official serial baseline via speculative decoding, with 54 promoted submissions from 14 solvers spanning GPT-6 Astra, DeepSeek-V4, Fable 5.1, Opus 5 and more. The author says MLX is close to a 100%+ improvement milestone, with CUDA grinding on.
More from Infra
- Nearly all top-10 PFAS makers plan production hikes to serve AI chips and data center cooling — jathansadowski · 2026-09-17
- Idle models might be smarter: a musing on GEMV vs GEMM under low concurrency — karminski3 · 2026-09-17
- University of Memphis study finds no major air quality deterioration around xAI's Colossus 1 — TinfoilTricorn · 2026-09-17
- Beam moves ~1TB every 30 minutes six months after launching on Bittensor — markjeffrey · 2026-09-17
- VC-Attention: training-free low-bit attention hits 1.9x on B200, beating FlashAttention-4 — xiuyu_l · 2026-09-17
- IBM NorthPole claims 22x inference performance over Nvidia on 12nm process — Site-Staff · 2026-09-17