Perplexity's Lily beats MLX-LM with 1.23x prefill and 1.35x decode throughput on M5 Max

perplexity_ai · x · 2026-09-03

Perplexity published benchmarks of Lily running Qwen3.6-35B-A3B on an M5 Max MacBook Pro:

Key idea: prefill and decode are fundamentally different workloads — prefill processes many prompt tokens at once with weight reuse, while decode generates one token at a time, making memory traffic and bandwidth the bottleneck.

Original post →

More from Infra

Infra channel →