Perplexity open-sources Lily, a local inference engine for Qwen3.6 on Apple silicon
perplexity_ai · x · 2026-09-03
Perplexity has open-sourced Lily, the local inference engine built for hybrid compute in Perplexity Computer.
- Specialized for Qwen3.6-35B-A3B on Apple silicon, so on-device compute doesn't bottleneck Computer tasks
- Designed for a hybrid setup splitting work between cloud models and a local model on the Mac
- Code available on GitHub (pplx-garden/lily)
More from Infra
- llama.cpp deprecates --chat-template-kwargs, reasoning-preserve now on by default — Bulky-Priority6824 · 2026-09-03
- Agentic API adds a stateful layer in front of vLLM for open-model agent runtimes — techNmak · 2026-09-03
- Google's Gemini 3.8 Flash 'works harder' but may burn more tokens at same pricing — The Verge AI · 2026-09-03
- Mitchell Hashimoto Details Memory Optimization Tricks in the Superlogical Server — sull · 2026-09-03
- Perplexity's Lily beats MLX-LM with 1.23x prefill and 1.35x decode throughput on M5 Max — perplexity_ai · 2026-09-03
- FastH3 Now Runs Locally on Apple Silicon and DGX Spark — Vandy_simp · 2026-09-03