Splash on a 40-core M5 Max: +20% decode speed by running tune-kernels on your own chip

SeveralViolins · reddit · 2026-09-25

A Reddit user benchmarked the Splash local inference engine on a 40-core M5 Max, where default kernel rules (tuned on smaller chips) underperform. Using Splash's built-in make tune-kernels tool, they found split-K layouts much faster for the 8-row draft-token step, gaining 20% decode speed. The post includes full build steps (source build, Xcode 26+ with Metal toolchain, tune-kernels with --confirm), an upstream issue (#154), and Swift model conversions on Hugging Face.

Original post →

More from Infra

Infra channel →