Perplexity open-sources Lily, a Rust/Metal inference server mapping Qwen ops to Apple silicon

perplexity_ai · x · 2026-09-03

Perplexity released and open-sourced Lily, designed to treat Apple silicon as a distinct inference platform and map Qwen's operations directly onto its compute and memory architecture.

GitHub repo (pplx-garden/lily, Rust) highlights:

The launch post also recaps the thread's benchmarks: 1.23x prefill and 1.35x decode throughput vs MLX-LM on M5 Max.

Related event: Perplexity Open-Sources Lily, a Local Inference Engine Built for Apple Silicon(6 posts)→

Original post →

More from Infra

Infra channel →