DeepGrove Launches Maple-Preview for On-Device Inference at 200+ Tokens/s
DeepGrove released Maple-Preview, an open-source 20B (1B activated) ternary-weight inference LLM. Achieving SOTA in its class, the model runs efficiently on edge devices like Mac Mini at over 200 tokens/s and can solve IMO-level math problems.
2026-08-05 ~ 2026-08-05 · 2 related posts
- 20B Model Maple-Preview Runs at 200+ tokens/s on Mac Mini, Solves IMO Math — tylerbruno05 · 2026-08-05
1 near-duplicate retellings: ycombinator