RTX 5090 Tested: Glimmer Model Hits 233.4 tps in Local Inference
YetAnotherAnonymoose · reddit · 2026-08-11
A Reddit user shared benchmarks running the Glimmer model with Dflash on the upcoming RTX 5090, achieving an astonishing 233.4 tps in local inference.
The poster noted the performance is incredibly promising, highlighting that a 256k context window is easily achievable on 24GB VRAM, outperforming Qwen in this regard. They plan to test it further on their own 4090.
Related event: RTX 5090 Hits 233 tps in Local Inference Test(2 posts)→
More from Infra
- Think Tank: AI Data Center Backlash Is Mostly Misguided and Fueled by Policy Failures — sebkrier · 2026-08-11
- llama.cpp Adds Cost-Based Tensor Split for 3-4% Speedup on Hybrid Multi-GPUs — milpster · 2026-08-11
- Agent Context Bottleneck: 99.9% Cache Hit Rate Masks Inefficiency — teortaxesTex · 2026-08-11
- OpenAI Seeks Power Trading Lead to Steer Global Data Center Energy Strategy — pstAsiatech · 2026-08-11
- Achieving 253 t/s on RTX 5090: Optimizing Muse Glimmer 30B Inference — patricious · 2026-08-11
- AI infrastructure investment hits 2.8% of US GDP, surpassing the railroad boom — QuintinPope5 · 2026-08-11