Open source doubles M5 Ultra MLX token prefill in just one week via Flash-Next

TheMoonMidas · x · 2026-09-29

Developer viticci benchmarked recent oMLX upstream builds and found MLX performance on the M5 Ultra has improved dramatically in a single week: token prefill has essentially doubled across the board thanks to Flash-Next. Maintainer jundotkim reshared the results, noting most of the gains came from open source contributors rather than himself, and thanked him for measuring properly. The tester says he'll need to update his review soon.

Original post →

More from Infra

Infra channel →