Maple 20B MoE Hits 9,885 tokens/s on a Single GH200, Preview Weights Released
MaziyarPanahi · x · 2026-08-06
DeepGroove has released the preview weights for the Maple 20B-A1B model on Hugging Face. It is a Mixture-of-Experts (MoE) model with 20 billion total parameters and 1 billion active parameters.
During testing, the model achieved an impressive throughput of 9,885 tokens/s on a single NVIDIA GH200 GPU with 64 concurrent requests. The author has opened a free endpoint for the community to stress-test and looks forward to the final post-trained model.
Related event: Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200(4 posts)→
More from Infra
- Developers Request 3-Tier MoE Offloading (Disk/CPU/GPU) for Local LLMs — storm1er · 2026-08-06
- AI Compute in Orbit: A Napkin Math Breakdown of SpaceX's Space Data Center — teortaxesTex · 2026-08-06
- Chamath Warns: AI Token Bill Doubles Every 45 Days While Productivity Grows Just 5% — rohanpaul_ai · 2026-08-06
- SanDisk projects NAND market revenue to exceed $300 billion in 2026 — Beth_Kindig · 2026-08-06
- Chips and Cheese Analyzes Nvidia's Vera Architecture Whitepaper — pella · 2026-08-06
- Local AI Hardware Guide: Choosing Between RTX 50-series and AMD for MiniMax H3 — Eden1506 · 2026-08-06