DeepGrove Releases On-Device LLM Running at 200+ tokens/s on Mac Mini
ycombinator · x · 2026-08-05
DeepGrove has released Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM.
Claimed to be SOTA in its weight class, the model can solve IMO-level problems. Most notably, it runs at 200+ tokens per second on a Mac Mini M4, making it 5–16× faster than efficient models of similar size like Gemma 4 and Qwen3.5, bringing always-on personal agents closer to reality.
Related event: DeepGrove Launches Maple-Preview for On-Device Inference at 200+ Tokens/s(2 posts)→
More from Infra
- Garage Solar-Powered 384GB Xeon Rig for Remote AI Coding — angadsg · 2026-08-05
- AMD CEO Lisa Su to Analyst: Your Data Center AI Number is Probably Too Low — firstadopter · 2026-08-05
- Open-Sourced Recipe: Running a 27B Local Agent 24/7 on a Single RTX 5090 — max_paperclips · 2026-08-05
- Europe Pledges €30B for AI Gigafactories, Only €1B Actually Committed — sanjaykalra · 2026-08-05
- Modal Optimizes Serverless Architecture to Reduce Network Latency — AAAzzam · 2026-08-05
- Elon Musk: AI Memory Demand Growing Over 200% Annually, Supply Lagging — firstadopter · 2026-08-05