Maple 20B MoE Hits 9,885 tokens/s on a Single GH200, Preview Weights Released

MaziyarPanahi · x · 2026-08-06

DeepGroove has released the preview weights for the Maple 20B-A1B model on Hugging Face. It is a Mixture-of-Experts (MoE) model with 20 billion total parameters and 1 billion active parameters.

During testing, the model achieved an impressive throughput of 9,885 tokens/s on a single NVIDIA GH200 GPU with 64 concurrent requests. The author has opened a free endpoint for the community to stress-test and looks forward to the final post-trained model.

Related event: Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200(4 posts)→

Original post →

More from Infra

Infra channel →