Muse Glimmer 30B Hits 3,323 tok/s on Single GH200, API Opened for Public Testing

MaziyarPanahi · x · 2026-08-12

Developer mkurman88 deployed the Muse Glimmer 30B model on a single NVIDIA GH200 GPU, achieving an astonishing inference speed of 3,323 tokens/s using the vLLM framework.

More interestingly, they have opened the API for public testing, inviting the community to send difficult prompts to find out where the model breaks.

Related event: Single GH200 GPU Achieves 3323 tok/s in Inference Test(2 posts)→

Original post →

More from Infra

Infra channel →