Perplexity open-sources Lily, a minimal Metal inference server for Qwen3.6-35B-A3B on M5 Macs
AravSrinivas · x · 2026-09-03
Perplexity CEO Arav Srinivas announced pplx-garden/lily, an open-source Metal inference server built for one checkpoint: Qwen3.6-35B-A3B converted to MLX affine 4-bit weights.
- Exposes a minimal OpenAI-compatible chat completions API; always decodes greedily
- Requires Apple GPU family 10+ (M5 or newer), macOS 26+, Rust 1.92
- Metal kernels compile from source at runtime, no offline shader build
- Performance reports include measurement contracts and repro steps; architecture and quantization layout are validated at load time
Related event: Perplexity Open-Sources Lily, a Local Inference Engine for Apple Silicon(7 posts)→
More from Infra
- GLM-5.3-Flash hits 1,005 tok/s locally on dual RTX PRO 6000 Blackwell cards — BanghuaZ · 2026-09-03
- Nvidia now ~8% of S&P 500 market cap, worth 16.3% of US GDP at $5.3T — ivan_bezdomny · 2026-09-03
- Broadcom Q3 AI chip revenue hits $16.7B, up 221% YoY; guides $21.7B for Q4 — Beth_Kindig · 2026-09-03
- PINNACLE: Claude Fable 5.1 halves agent failure rate but costs $2.46 per correct answer — ryanshrout · 2026-09-03
- eBPF looks cheap but isn't free: Bitbison's deep dive on hooks, CO-RE reads and rings — tianyin_xu · 2026-09-03
- VoxGen: a Rust + Vulkan TTS engine bringing AMD support to VoxCPM2 without PyTorch — Substantial_Swan_144 · 2026-09-03