DeepGrove Releases On-Device LLM Running at 200+ tokens/s on Mac Mini

ycombinator · x · 2026-08-05

DeepGrove has released Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM.

Claimed to be SOTA in its weight class, the model can solve IMO-level problems. Most notably, it runs at 200+ tokens per second on a Mac Mini M4, making it 5–16× faster than efficient models of similar size like Gemma 4 and Qwen3.5, bringing always-on personal agents closer to reality.

Related event: DeepGrove Launches Maple-Preview for On-Device Inference at 200+ Tokens/s(2 posts)→

Original post →

More from Infra

Infra channel →