Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests
MaziyarPanahi · x · 2026-08-06
A developer benchmarked DeepGroove AI's new Maple 20B-A1M model on a single NVIDIA GH200, achieving a throughput of 9,885 tokens/s with 64 concurrent requests.
To evaluate its performance for on-device tasks, the author opened a temporary, free API endpoint for 12 hours, encouraging the community to stress-test it and build applications.
Related event: Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200(4 posts)→
More from Infra
- Chamath Warns: AI Token Bill Doubles Every 45 Days While Productivity Grows Just 5% — rohanpaul_ai · 2026-08-06
- SanDisk projects NAND market revenue to exceed $300 billion in 2026 — Beth_Kindig · 2026-08-06
- Local AI Hardware Guide: Choosing Between RTX 50-series and AMD for MiniMax H3 — Eden1506 · 2026-08-06
- Samsung to Lock 60-70% of Production in Long-Term Deals, Tech Giants as Key Clients — Beth_Kindig · 2026-08-06
- SanDisk Executives Assert: Over 80% Gross Margin is a 'Fair Return' — firstadopter · 2026-08-06
- Google's AI token processing surges 330x in two years, signaling booming inference demand — Beth_Kindig · 2026-08-06