Perplexity explains Lily: purpose-built inference vs general-purpose MLX-LM framework

perplexity_ai · x · 2026-09-03

Perplexity elaborated on Lily's positioning: hybrid compute splits work between cloud models and a local model on the Mac, so local inference must keep pace with the rest of the task. MLX-LM is a general-purpose framework, while Lily is purpose-built for this inference workload — treating prefill (many tokens at once, weight reuse) and decode (one token at a time, memory-bandwidth-bound) as fundamentally different.

Related event: Perplexity Open-Sources Lily, a Local Inference Engine Built for Apple Silicon(6 posts)→

Original post →

More from Infra

Infra channel →