Perplexity open-sources Lily, a Metal inference engine for Qwen on Apple Silicon, 35% faster decode than MLX-LM

inductionheads · x · 2026-09-03

Perplexity open-sourced Lily, the local inference engine built for hybrid compute in Perplexity Computer, specialized for Qwen3.6-35B-A3B (MLX affine 4-bit) on Apple Silicon.

Key details:

Exposes a minimal OpenAI chat completions subset, greedy decoding only; requires Apple GPU family 10+ (M5 or newer) and macOS 26+. Code, blog, and reproduction steps are public.

Related event: Perplexity Open-Sources Lily, a Qwen Inference Engine Built for Apple Silicon(5 posts)→

Original post →

More from Infra

Infra channel →