lithos-metal Megakernels Beat Ollama on M5 Pro, Tuned for M5 Max

JiaZhihao · x · 2026-10-09

lithos-metal author JiaZhihao confirmed its megakernels are heavily tuned for M5 Max with M5 Pro tuning underway. A user's real-world test running Qwen3 27B on M5 Pro showed lithos meaningfully faster than Ollama after warmup — though the voice-IME workload (5.6K-token system prompt, 15-90 output tokens, greedy decoding) cares about latency and model keep-hot/idle-unload UX more than throughput, where Ollama still leads on features.

Original post →

More from Infra

Infra channel →