Perplexity details why hybrid compute needs a purpose-built local engine like Lily
perplexity_ai · x · 2026-09-03
Perplexity's thread recaps why it open-sourced Lily: built for hybrid compute in Perplexity Computer, where cloud models and a local Mac model split the work and on-device inference must not bottleneck tasks. Lily is specialized for Qwen3.6-35B-A3B on Apple silicon, treating prefill and decode as distinct workloads unlike the general-purpose MLX-LM.
More from Infra
- Broadcom guides AI revenue to ~$115B in FY2027, doubling again to $230B in FY2028 — BenBajarin · 2026-09-03
- Perplexity open-sources Lily, its Apple Silicon inference engine for Qwen3.6-35B-A3B — inductionheads · 2026-09-03
- Nvidia now ~8% of S&P 500 market cap, worth 16.3% of US GDP at $5.3T — ivan_bezdomny · 2026-09-03
- Broadcom Q3 AI chip revenue hits $16.7B, up 221% YoY; guides $21.7B for Q4 — Beth_Kindig · 2026-09-03
- PINNACLE: Claude Fable 5.1 halves agent failure rate but costs $2.46 per correct answer — ryanshrout · 2026-09-03
- eBPF looks cheap but isn't free: Bitbison's deep dive on hooks, CO-RE reads and rings — tianyin_xu · 2026-09-03