mlx.fast fixes speed-display bug: MLX kernels hit 80.6 tps, nearing 100% speedup milestone

HankYeomans · x · 2026-09-17

TheDavidTai, author of the mlx.fast inference-optimization project, fixed two bugs:

The leaderboard's current record stands at a 97.3% composite speedup over the official serial baseline via speculative decoding, with 54 promoted submissions from 14 solvers spanning GPT-6 Astra, DeepSeek-V4, Fable 5.1, Opus 5 and more. The author says MLX is close to a 100%+ improvement milestone, with CUDA grinding on.

Original post →

More from Infra

Infra channel →