A GLM-5.2 inference debate asks how 750B parameters can exceed 1 token per second

francoisfleuret · x · 2026-07-21

A quoted thread asks how GLM-5.2 can run at better than 1 token/s when it is said to have 750B parameters and 40B active parameters—roughly 20 GB—while even a top SSD only reaches about 15 GB/s in theory.

The author of the post says they wish they could summarize the responses they received and correct the mistake in the original claim, implying the discussion is about reconciling the model’s apparent size with the actual inference bottleneck.

Original post →

More from Infra

Infra channel →