RTX 5090 runs 30B model inference at over 233 tps in tests
Benchmark tests show the upcoming RTX 5090 achieves over 233 tokens per second running the 30B Muse Glimmer model, with speculative decoding vastly outperforming Mac.
2026-08-10 ~ 2026-08-11 · 3 related posts
- Speculative Decoding Boosts RTX 5090 to 233 tok/s, Outperforming Mac — rohanpaul_ai · 2026-08-10
- RTX 5090 Tested: Glimmer Model Hits 233.4 tps in Local Inference — YetAnotherAnonymoose · 2026-08-11
- Achieving 253 t/s on RTX 5090: Optimizing Muse Glimmer 30B Inference — patricious · 2026-08-11