FULL STORY
DeepGrove Maple: From Launch to Extreme Speed Tests
DeepGrove released the open-source Maple-Preview model, which was subsequently tested by developers demonstrating extreme inference speeds on both Mac Mini and a single NVIDIA GH200.
2026-08-05 ~ 2026-08-06 · 2 episodes · 8 posts
Episode 1 · DeepGrove's open-source Maple model runs at 200+ tokens/s on Mac Mini (2026-08-05, 4 posts)
DeepGrove released Maple-Preview, an open-source 20B parameter ternary-weight LLM. Achieving SOTA performance in its class, the model tackles complex math and runs at over 200 tokens per second on consumer hardware like a Mac Mini.
- 20B Model Maple-Preview Runs at 200+ tokens/s on Mac Mini, Solves IMO Math — tylerbruno05 · 2026-08-05
- DeepGrove Releases On-Device LLM Running at 200+ tokens/s on Mac Mini — ycombinator · 2026-08-05
- Maple-Preview: Open-Source Ternary-Weight LLM Hits 200+ tokens/s on Mac Mini — garrytan · 2026-08-05
- Deepgrove Releases Maple-Preview: 20B Ternary-Weight Open-Weight LLM — cafedude · 2026-08-05
Episode 2 · Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200 (2026-08-06, 4 posts)
DeepGroove released a preview of the open-source MoE model Maple 20B-A1B. Tests show it achieves an impressive 9,885 tokens/s throughput on a single NVIDIA GH200 processor with 64 concurrent requests. A free inference API endpoint has been temporarily opened for community testing.
- Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests — MaziyarPanahi · 2026-08-06
- Maple 20B Hits 9.8k Tokens/s on Single GH200, Opens Free Inference Endpoint — MaziyarPanahi · 2026-08-06
- Maple 20B MoE Hits 9,885 tokens/s on a Single GH200, Preview Weights Released — MaziyarPanahi · 2026-08-06
- Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests — MaziyarPanahi · 2026-08-06