Perplexity's Lily beats MLX-LM with 1.23x prefill and 1.35x decode throughput on M5 Max
perplexity_ai · x · 2026-09-03
Perplexity published benchmarks of Lily running Qwen3.6-35B-A3B on an M5 Max MacBook Pro:
- Across ten prompt lengths and ten decode contexts, Lily was consistently faster than MLX-LM
- On average 1.23x higher prefill throughput and 1.35x higher decode throughput
- Output quality effectively unchanged
Key idea: prefill and decode are fundamentally different workloads — prefill processes many prompt tokens at once with weight reuse, while decode generates one token at a time, making memory traffic and bandwidth the bottleneck.
More from Infra
- llama.cpp deprecates --chat-template-kwargs, reasoning-preserve now on by default — Bulky-Priority6824 · 2026-09-03
- Agentic API adds a stateful layer in front of vLLM for open-model agent runtimes — techNmak · 2026-09-03
- Google's Gemini 3.8 Flash 'works harder' but may burn more tokens at same pricing — The Verge AI · 2026-09-03
- Mitchell Hashimoto Details Memory Optimization Tricks in the Superlogical Server — sull · 2026-09-03
- Perplexity open-sources Lily, a local inference engine for Qwen3.6 on Apple silicon — perplexity_ai · 2026-09-03
- FastH3 Now Runs Locally on Apple Silicon and DGX Spark — Vandy_simp · 2026-09-03