Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests

MaziyarPanahi · x · 2026-08-06

Developer @MaziyarPanahi benchmarked the newly released open-source model Maple 20B-A1B by @deepgroveai. On a single NVIDIA GH200 processor with 64 concurrent requests, the model achieved a staggering inference speed of 9,885 tokens/s.

The author subsequently opened a free testing endpoint for the model (limited to 12 hours), inviting the community to stress-test it. The model is also highly anticipated to become a medical AI application capable of running locally on smartphones.

Related event: Open-Source MoE Model Maple 20B Hits Record Inference Speeds on Single GH200(2 posts)→

Original post →

More from Infra

Infra channel →