Perplexity explains Lily: purpose-built inference vs general-purpose MLX-LM framework
perplexity_ai · x · 2026-09-03
Perplexity elaborated on Lily's positioning: hybrid compute splits work between cloud models and a local model on the Mac, so local inference must keep pace with the rest of the task. MLX-LM is a general-purpose framework, while Lily is purpose-built for this inference workload — treating prefill (many tokens at once, weight reuse) and decode (one token at a time, memory-bandwidth-bound) as fundamentally different.
More from Infra
- Broadcom guides AI revenue to ~$115B in FY2027, doubling again to $230B in FY2028 — BenBajarin · 2026-09-03
- Perplexity open-sources Lily, its Apple Silicon inference engine for Qwen3.6-35B-A3B — inductionheads · 2026-09-03
- Nvidia now ~8% of S&P 500 market cap, worth 16.3% of US GDP at $5.3T — ivan_bezdomny · 2026-09-03
- Broadcom Q3 AI chip revenue hits $16.7B, up 221% YoY; guides $21.7B for Q4 — Beth_Kindig · 2026-09-03
- PINNACLE: Claude Fable 5.1 halves agent failure rate but costs $2.46 per correct answer — ryanshrout · 2026-09-03
- eBPF looks cheap but isn't free: Bitbison's deep dive on hooks, CO-RE reads and rings — tianyin_xu · 2026-09-03