Perplexity Mac Hybrid mode: Local models handle subtasks
testingcatalog · x · 2026-08-31
Perplexity is launching a Hybrid mode for Mac, allowing local models to handle specific subtasks to save costs. Supported models include Gemma 4 (16GB RAM), Qwen 2.5 32B (32GB RAM), and a proprietary Perplexity model (32GB RAM). Cloud orchestrates heavy reasoning while local inference handles lighter tasks. A new 'Privacy Gate' feature runs locally to inspect data for PII before cloud transmission.
Related event: Perplexity Brings Hybrid Mode to Mac With Local Models(2 posts)→
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- 2 engineers + AI designed a working LLM chip in 2 weeks, no human in the loop — 新智元 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01