Open-source lithos-metal hits 200+ tokens/s/user on Qwen3.8-27B with one M5 Max

JiaZhihao · x · 2026-10-09

lithos-metal is now fully open source: an inference engine that generates Metal megakernels for Apple silicon, combining layer-wise mixer megakernels, GPU-resident generation control, and DSpark speculative decoding. It runs Qwen3.8-27B at 200+ tokens/s/user peak on a single M5 Max, with local serving that plugs into any coding agent in one command. Apache-2.0 licensed, code and tech blog available.

Original post →

More from coding & agent

coding & agent channel →