Perplexity details why hybrid compute needs a purpose-built local engine like Lily

perplexity_ai · x · 2026-09-03

Perplexity's thread recaps why it open-sourced Lily: built for hybrid compute in Perplexity Computer, where cloud models and a local Mac model split the work and on-device inference must not bottleneck tasks. Lily is specialized for Qwen3.6-35B-A3B on Apple silicon, treating prefill and decode as distinct workloads unlike the general-purpose MLX-LM.

Related event: Perplexity Open-Sources Lily, a Local Inference Engine Built for Apple Silicon(6 posts)→

Original post →

More from Infra

Infra channel →