Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200

DeepGroove released a preview of the open-source MoE model Maple 20B-A1B. Tests show it achieves an impressive 9,885 tokens/s throughput on a single NVIDIA GH200 processor with 64 concurrent requests. A free inference API endpoint has been temporarily opened for community testing.

2026-08-06 ~ 2026-08-06 · 4 related posts

Full story(2 episodes)→

1 near-duplicate retellings: MaziyarPanahi