Maple 20B Hits 9,885 tokens/s on a Single NVIDIA GH200 in Concurrency Test
MaziyarPanahi · x · 2026-08-06
Developer MaziyarPanahi benchmarked the open-source Maple 20B-A1B model developed by DeepGrooveai.
- Hardware & Performance: Running on a single NVIDIA GH200 with 64 concurrent requests, the model achieved an impressive throughput of 9,885 tokens/s.
- Open Testing: The author briefly opened a free inference endpoint for the community to stress-test, validating its stability under heavy concurrent streaming workloads.
Related event: Maple 20B Open-Source MoE Hits 9,885 tokens/s on Single GH200(5 posts)→
More from Infra
- PlanetScale Details Massively Parallel Backups: Restoring PB-Scale Databases at 50GB/s — bibryam · 2026-08-06
- Testing MiniMax H3 native ComfyUI on RTX 3060 12GB: Setup guide — Creepy-Fault6977 · 2026-08-06
- Calling for Decentralized AI: From the Cloud Back to a $5K Local Personalized AGI — Dan_Jeffries1 · 2026-08-06
- Running 28.9M LLM on $8 ESP32: 10 tok/s Fully On-Device — FinanceYF5 · 2026-08-06
- Running Minimax Video on 12GB VRAM: 22-Min Render Tested — Svan_Derh · 2026-08-06
- Sergey Brin: Jeff Dean Pushed for Google's TPUs After 3-Minute Voice Usage Doubled CPU Demand — rohanpaul_ai · 2026-08-06