MLX MoE Layer Gets 1.5x Faster via Better Tile Scheduling in Grouped Matmul
awnihannun · x · 2026-09-28
A merged MLX PR reworks the gathermm kernel's Metal tile scheduling: previously tiles were assigned to thread groups independently of expert assignment, forcing extra matmul passes and redundant K scans (worst on short sequences). The new approach, inspired by quack's tile scheduler, aligns tile loading with expert assignment, delivering 1.5x speedup for MoE layers on Apple Silicon.
More from Infra
- Lightning AI and Google Cloud cut PyTorch Lightning checkpoint write times by up to 95% — LightningAI · 2026-09-28
- Crucible Capital founder financed her own GPU cluster with stablecoin debt and personal credit risk — MarvinTBaumann · 2026-09-28
- NVIDIA launches Open Agent Safety Platform for controlling what AI agents can do — nvidia · 2026-09-28
- Cloudflare incident: skipped block zeroing leaked tenant data across 18 of 24 containers — arpit_bhayani · 2026-09-28
- On RTX 5090, Qwen 27B hits 200 TPS but Flash next only 50: what model sits between for coding? — MasterNomie · 2026-09-28
- Lumen Launches On-Demand Dedicated Internet Up to 100 Gbps at 10M US Sites — shashib · 2026-09-28