Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200
DeepGroove released a preview of the open-source MoE model Maple 20B-A1B. Tests show it achieves an impressive 9,885 tokens/s throughput on a single NVIDIA GH200 processor with 64 concurrent requests. A free inference API endpoint has been temporarily opened for community testing.
2026-08-06 ~ 2026-08-06 · 4 related posts
- Episode 1: DeepGrove's open-source Maple model runs at 200+ tokens/s on Mac Mini(2026-08-05, 4 posts)
- Episode 2: Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200(2026-08-06, 4 posts)
- Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests — MaziyarPanahi · 2026-08-06
- Maple 20B Hits 9.8k Tokens/s on Single GH200, Opens Free Inference Endpoint — MaziyarPanahi · 2026-08-06
- Maple 20B MoE Hits 9,885 tokens/s on a Single GH200, Preview Weights Released — MaziyarPanahi · 2026-08-06
1 near-duplicate retellings: MaziyarPanahi