Perplexity open-sources Lily, a Rust/Metal inference server mapping Qwen ops to Apple silicon
perplexity_ai · x · 2026-09-03
Perplexity released and open-sourced Lily, designed to treat Apple silicon as a distinct inference platform and map Qwen's operations directly onto its compute and memory architecture.
GitHub repo (pplx-garden/lily, Rust) highlights:
- A Metal inference server for one checkpoint: Qwen3.6-35B-A3B as MLX affine 4-bit weights (group size 64)
- Exposes a minimal subset of the OpenAI chat completions API, always decodes greedily
- Metal kernels compile from source at runtime, no offline shader build
- Requires Apple GPU family 10+ (M5 and newer), macOS 26+, Rust 1.92
- Validates exact architecture and quantization layout at load time, with benchmark reports and reproduction steps
The launch post also recaps the thread's benchmarks: 1.23x prefill and 1.35x decode throughput vs MLX-LM on M5 Max.
More from Infra
- Broadcom guides AI revenue to ~$115B in FY2027, doubling again to $230B in FY2028 — BenBajarin · 2026-09-03
- Perplexity open-sources Lily, its Apple Silicon inference engine for Qwen3.6-35B-A3B — inductionheads · 2026-09-03
- Nvidia now ~8% of S&P 500 market cap, worth 16.3% of US GDP at $5.3T — ivan_bezdomny · 2026-09-03
- Broadcom Q3 AI chip revenue hits $16.7B, up 221% YoY; guides $21.7B for Q4 — Beth_Kindig · 2026-09-03
- PINNACLE: Claude Fable 5.1 halves agent failure rate but costs $2.46 per correct answer — ryanshrout · 2026-09-03
- eBPF looks cheap but isn't free: Bitbison's deep dive on hooks, CO-RE reads and rings — tianyin_xu · 2026-09-03