Maple 20B Hits 9.8k Tokens/s on Single GH200, Opens Free Inference Endpoint
MaziyarPanahi · x · 2026-08-06
Developer Maziyar Panahi announced the opening of a free inference API endpoint for the Maple 20B-A1B model to the community for a 12-hour limited time.
Deployed on a single NVIDIA GH200 using vLLM 0.26.0 and FP8 precision, the endpoint achieved an impressive throughput of 9,885 tokens/s with 64 concurrent requests. Users can test and stress it using the provided address and key, restricted to synthetic or public data only.
Related event: Maple 20B Open-Source Model Hits 9,885 tokens/s on Single GH200(4 posts)→
More from Infra
- Samsung to Lock 60-70% of Production in Long-Term Deals, Tech Giants as Key Clients — Beth_Kindig · 2026-08-06
- SanDisk Executives Assert: Over 80% Gross Margin is a 'Fair Return' — firstadopter · 2026-08-06
- Google's AI token processing surges 330x in two years, signaling booming inference demand — Beth_Kindig · 2026-08-06
- Long Contexts Multiply Speculative Decoding Gains, Acceptance Rate Nears 100% — DjCanalex · 2026-08-06
- The Enterprise AI Question: Where Does Your AI Actually Run? — DavidLinthicum · 2026-08-06
- AI Inference Demand Growing 10x Yearly Will Make Compute Scarcity the Default — TansuYegen · 2026-08-06