Developer criticizes MLX for low bandwidth utilization on M5 Max
andrejusb · x · 2026-08-26
Developer @Youssofal criticized Apple's MLX framework for poor performance, noting that BF16 training only saturates 220-230GB/s on the M5 Max (614GB/s bandwidth), wasting resources. He blamed immature kernels rather than hardware. @ivanfioravanti countered that such arrogant tones are unhelpful and emphasized the need for collaboration over conflict.
More from Infra
- Free Qwen3.8-Flash-Next endpoint launched on 4× H200 at 100+ tok/s — victormustar · 2026-08-26
- Signadot Enables Second-Scale PR Preview Environments — pjausovec · 2026-08-26
- 1-bit residuals shrink index 13x for 1.6 point loss — tomaarsen · 2026-08-26
- Optimized LI index smaller, beats Qwen3-8B by 9 points — tomaarsen · 2026-08-26
- Dev Mines 36B Free Tokens from Zhipu for Open Source Projects — doodlestein · 2026-08-26
- Edge sensor nodes struggle to disconnect from hyper-centralized foundation models — curious_vii · 2026-08-26