GLM-5.2 inference on RTX 5090s jumps from 30 tok/s to 80–110 tok/s

markjeffrey · x · 2026-07-27

GLM-5.2 inference on RTX 5090s is now 3× faster

A post quoting Ning says GLM-5.2 inference on RTX 5090s has jumped from roughly 30 tok/s to 80–110 tok/s on a single stream — about 3× faster overnight.

The update is framed as a live performance improvement for local inference, not a new model release.

Original post →

More from Infra

Infra channel →