Muse Glimmer 30B Hits 3,323 tok/s on Single GH200, API Open for Testing

MaziyarPanahi · x · 2026-08-12

A developer has deployed the Muse Glimmer 30B model on a single NVIDIA GH200 GPU using DFlash, achieving an impressive inference speed of 3,323 tokens/s. The author has opened the API for public testing, inviting the community to push the model to its limits with difficult prompts. Additionally, a similar test for the Qwen3.8 27B model is scheduled for tomorrow.

Related event: Single GH200 Hits 3323 tok/s in Muse Glimmer 30B Inference(3 posts)→

Original post →

More from Infra

Infra channel →