Maple-Preview: Open-Source Ternary-Weight LLM Hits 200+ tokens/s on Mac Mini

garrytan · x · 2026-08-05

DeepGrove has introduced Maple-Preview, an open-source 20B parameter (1B active) ternary-weight reasoning LLM.

The model achieves state-of-the-art performance within its weight class and is capable of solving IMO-level math problems. Thanks to its ternary weights, it runs at over 200 tokens per second on a Mac Mini M4, making it 5–16× faster than comparable efficient models like Gemma and Qwen.

Related event: DeepGrove's open-source Maple model runs at 200+ tokens/s on Mac Mini(4 posts)→

Original post →

More from Infra

Infra channel →