GLM-5.2 inference on RTX 5090s jumps from 30 tok/s to 80–110 tok/s
markjeffrey · x · 2026-07-27
GLM-5.2 inference on RTX 5090s is now 3× faster
A post quoting Ning says GLM-5.2 inference on RTX 5090s has jumped from roughly 30 tok/s to 80–110 tok/s on a single stream — about 3× faster overnight.
The update is framed as a live performance improvement for local inference, not a new model release.
More from Infra
- Voice AI’s real call cost is more than minutes: STT, TTS, SIP and retries — decant338 · 2026-07-27
- Micron and Meta paper says Spark can slow down 38× when shuffle spills to SSD — dr_alphalyrae · 2026-07-27
- Inference businesses may be charging 7× to 15× more than renting a GPU — JoshPurtell · 2026-07-27
- CUDA benchmark shows 64M-number sum runs 20× faster on GPU than CPU — ctjlewis · 2026-07-27
- AI buildout pushes data-center financing into more creative capital structures — GaryMarcus · 2026-07-27
- Users ask whether RAID 0 NVMe setups improve local large-model runs on Pulsar — Wyldkard79 · 2026-07-27