Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests

MaziyarPanahi · x · 2026-08-06

A developer benchmarked DeepGroove AI's new Maple 20B-A1M model on a single NVIDIA GH200, achieving a throughput of 9,885 tokens/s with 64 concurrent requests.

To evaluate its performance for on-device tasks, the author opened a temporary, free API endpoint for 12 hours, encouraging the community to stress-test it and build applications.

Related event: Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200(4 posts)→

Original post →

More from Infra

Infra channel →