Splash on a 40-core M5 Max: +20% decode speed by running tune-kernels on your own chip
SeveralViolins · reddit · 2026-09-25
A Reddit user benchmarked the Splash local inference engine on a 40-core M5 Max, where default kernel rules (tuned on smaller chips) underperform. Using Splash's built-in make tune-kernels tool, they found split-K layouts much faster for the 8-row draft-token step, gaining 20% decode speed. The post includes full build steps (source build, Xcode 26+ with Metal toolchain, tune-kernels with --confirm), an upstream issue (#154), and Swift model conversions on Hugging Face.
More from Infra
- Jensen Huang Pushes Back on AI Doom: "AI Is Software, It's Math" — DavidLinthicum · 2026-09-25
- Elon Musk: space compute will obviously round up to 100% of all compute — CurieuxExplorer · 2026-09-25
- Tessara's falsification test for its Micron call: watch DRAM price change on Sept 30 — tengyanAI · 2026-09-25
- Tessara predicts Micron Q3 revenue of $56.2B, beating the highest of 22 analyst estimates — tengyanAI · 2026-09-25
- Your data stack is about to get less forgiving: agents turn stale data into wrong actions — bigdata · 2026-09-25
- Bending Spoons runs 99% of AI traffic on self-hosted open models, thanks to evals — alex_verem · 2026-09-25