A GLM-5.2 inference debate asks how 750B parameters can exceed 1 token per second
francoisfleuret · x · 2026-07-21
A quoted thread asks how GLM-5.2 can run at better than 1 token/s when it is said to have 750B parameters and 40B active parameters—roughly 20 GB—while even a top SSD only reaches about 15 GB/s in theory.
The author of the post says they wish they could summarize the responses they received and correct the mistake in the original claim, implying the discussion is about reconciling the model’s apparent size with the actual inference bottleneck.
More from Infra
- Tesla’s FSD v14 Lite is reportedly headed to 4 million older HW3 cars — MatthewBerman · 2026-07-21
- TSMC’s 3nm utilization reportedly tops 120% as AI demand drives a $190B capex cycle — tengyanAI · 2026-07-21
- Nativ brings local AI model running to Mac with a desktop app and localhost API — Simon Willison · 2026-07-21
- Octen says agent search now runs at 62ms P50 with only a 6ms P90 gap — aakashgupta · 2026-07-21
- Zhipu acquires a compiler-team spinout to optimize AI inference on domestic chips — zephyr_z9 · 2026-07-21
- Open reproduction of Meta’s REWIRE data pipeline cuts the cost to about $11 — vanstriendaniel · 2026-07-21