lithos-metal Megakernels Beat Ollama on M5 Pro, Tuned for M5 Max
JiaZhihao · x · 2026-10-09
lithos-metal author JiaZhihao confirmed its megakernels are heavily tuned for M5 Max with M5 Pro tuning underway. A user's real-world test running Qwen3 27B on M5 Pro showed lithos meaningfully faster than Ollama after warmup — though the voice-IME workload (5.6K-token system prompt, 15-90 output tokens, greedy decoding) cares about latency and model keep-hot/idle-unload UX more than throughput, where Ollama still leads on features.
More from Infra
- Nunchux inference engine joins AMD's AI Inference Engines & Services ecosystem — junyanz89 · 2026-10-09
- tinygrad runs its GitHub Actions CI on 4 new tinybox machines — AIFlow_ML · 2026-10-09
- OpenAI's 10,000-agent, 130B-token run pushed slime v0.4.0 to rethink RL infrastructure scale — teortaxesTex · 2026-10-09
- TokenRouter Serving System Boosts Token-Level LLM Routing Throughput up to 64x — nics-efc · 2026-10-09
- SparseDecoding: Decoding-Aware Pruning Yields up to 1.48x Faster LLM Inference — encodelab · 2026-10-09
- x86_64 Edge runs with GPU acceleration on Arduino VENTUNO Q via FEX — unixterminal · 2026-10-09