Achieving 253 t/s on RTX 5090: Optimizing Muse Glimmer 30B Inference
patricious · reddit · 2026-08-11
The author benchmarked the Muse Glimmer 30B model (Q5KM quantization) on an RTX 5090 (32GB). By applying PR #26842, which moves the DFlash draft argmax from CPU to GPU, they successfully resolved a key bottleneck.
With this optimization, inference speed for code generation jumped from 78 t/s to 220-253 t/s, and mixed agent tasks reached 188-213 t/s. The post includes detailed launch flags and notes that reasoning budget flags currently do not work with this template.
Related event: RTX 5090 runs 30B model inference at over 233 tps in tests(3 posts)→
More from Infra
- Under 10% of Enterprises Scale AI; Compute Shortage to Persist — BenBajarin · 2026-08-11
- Amazon Backs Texas Gas Plant That May Become Top US Climate Polluter for AI — Ars Technica AI · 2026-08-11
- HKU Team Bypasses Von Neumann Bottleneck with Ultra-low Power 2D Material AI Chip — YiMaTweets · 2026-08-11
- Blockstream launches atomic swaps for Bitcoin and Lightning, citing AI-assisted attacks — RSync25 · 2026-08-11
- Local AI Video on RTX 3060 Ti: Generate High-Quality Clips in 20 Minutes with 8GB VRAM — cocktailpeanut · 2026-08-11
- Running 35B Models on a Single RTX 5080: Migrating from llama.cpp to vLLM — McFlurriez · 2026-08-11