RTX 5090 runs 30B model inference at over 233 tps in tests

Benchmark tests show the upcoming RTX 5090 achieves over 233 tokens per second running the 30B Muse Glimmer model, with speculative decoding vastly outperforming Mac.

2026-08-10 ~ 2026-08-11 · 3 related posts