Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests
MaziyarPanahi · x · 2026-08-06
Developer @MaziyarPanahi benchmarked the newly released open-source model Maple 20B-A1B by @deepgroveai. On a single NVIDIA GH200 processor with 64 concurrent requests, the model achieved a staggering inference speed of 9,885 tokens/s.
The author subsequently opened a free testing endpoint for the model (limited to 12 hours), inviting the community to stress-test it. The model is also highly anticipated to become a medical AI application capable of running locally on smartphones.
More from Infra
- The Enterprise AI Question: Where Does Your AI Actually Run? — DavidLinthicum · 2026-08-06
- AI Inference Demand Growing 10x Yearly Will Make Compute Scarcity the Default — TansuYegen · 2026-08-06
- If Your Data Can't Move, Your AI Strategy Is Doomed — DavidLinthicum · 2026-08-06
- Nativ brings LiquidAI LFM2.5 to Mac locally: 82 tok/s decode, <8.5GB RAM for 128K context — JosephJacks_ · 2026-08-06
- Startup Panthalassa Develops Floating AI Data Centers Powered by Wave Energy — TinfoilTricorn · 2026-08-06
- Maple 20B Hits 9.8k Tokens/s on Single GH200, Opens Free Inference Endpoint — MaziyarPanahi · 2026-08-06