Maple 20B Hits 9.8k Tokens/s on Single GH200, Opens Free Inference Endpoint

MaziyarPanahi · x · 2026-08-06

Developer Maziyar Panahi announced the opening of a free inference API endpoint for the Maple 20B-A1B model to the community for a 12-hour limited time.

Deployed on a single NVIDIA GH200 using vLLM 0.26.0 and FP8 precision, the endpoint achieved an impressive throughput of 9,885 tokens/s with 64 concurrent requests. Users can test and stress it using the provided address and key, restricted to synthetic or public data only.

Related event: Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200(4 posts)→

Original post →

More from Infra

Infra channel →