RTX 5090 Tested: Glimmer Model Hits 233.4 tps in Local Inference

YetAnotherAnonymoose · reddit · 2026-08-11

A Reddit user shared benchmarks running the Glimmer model with Dflash on the upcoming RTX 5090, achieving an astonishing 233.4 tps in local inference.

The poster noted the performance is incredibly promising, highlighting that a 256k context window is easily achievable on 24GB VRAM, outperforming Qwen in this regard. They plan to test it further on their own 4090.

Related event: RTX 5090 Hits 233 tps in Local Inference Test(2 posts)→

Original post →

More from Infra

Infra channel →