MLX MoE Layer Gets 1.5x Faster via Better Tile Scheduling in Grouped Matmul

awnihannun · x · 2026-09-28

A merged MLX PR reworks the gathermm kernel's Metal tile scheduling: previously tiles were assigned to thread groups independently of expert assignment, forcing extra matmul passes and redundant K scans (worst on short sequences). The new approach, inspired by quack's tile scheduler, aligns tile loading with expert assignment, delivering 1.5x speedup for MoE layers on Apple Silicon.

Original post →

More from Infra

Infra channel →