DeepGrove Launches Maple-Preview for On-Device Inference at 200+ Tokens/s

DeepGrove released Maple-Preview, an open-source 20B (1B activated) ternary-weight inference LLM. Achieving SOTA in its class, the model runs efficiently on edge devices like Mac Mini at over 200 tokens/s and can solve IMO-level math problems.

2026-08-05 ~ 2026-08-05 · 2 related posts

1 near-duplicate retellings: ycombinator