Open source doubles M5 Ultra MLX token prefill in just one week via Flash-Next
TheMoonMidas · x · 2026-09-29
Developer viticci benchmarked recent oMLX upstream builds and found MLX performance on the M5 Ultra has improved dramatically in a single week: token prefill has essentially doubled across the board thanks to Flash-Next. Maintainer jundotkim reshared the results, noting most of the gains came from open source contributors rather than himself, and thanked him for measuring properly. The tester says he'll need to update his review soon.
More from Infra
- More Americans oppose a local data center than a nuclear reactor, says Cathie Wood — PeterDiamandis · 2026-09-29
- Ollama 0.40 RC Tested: MLX Speed Is Real, But the Stable Version Already Has It — TheOyinbooke · 2026-09-29
- H3 T2VA VRAM squeeze sparks proposal for a remote text-encoder API to keep the 32B off your GPU — frankmanbb · 2026-09-29
- 30,000-word report: DUVi lithography exports will decide US-China AI chip race over the next decade — fiiiiiist · 2026-09-29
- xLLM open-sourced: flexible pre-training infra hits 10,050 tokens/s/GPU on H200 without dataset rebuilds — HongyiWang10 · 2026-09-29
- AMD Hits $1T Market Cap With Just 5%-7% of GPU Server Market — and the Author Says It's Still Undervalued — Beth_Kindig · 2026-09-29