Perplexity Open-Sources Lily, a Qwen Inference Engine Built for Apple Silicon

Perplexity announced in a tweet thread on September 3 the open-sourcing of Lily, a local inference engine built for the hybrid compute architecture of its Perplexity Computer. The code (pplx-garden/lily, implemented in Rust) is now live on GitHub. Lily is purpose-built for Qwen3.6-35B-A3B on Apple Silicon (MLX affine 4-bit weights), and a repost mentions a 35% boost in decoding speed.

Confirmed

Why it matters

2026-09-03 ~ 2026-09-03 · 5 related posts

Primary sources