GLM CPU-only Inference Questioned as Speed Falls Below 0.2 tok/s
Despite claims of running entirely on CPU, tests of the GLM 5.2 Colibri int4 model on an AMD5 machine yielded less than 0.2 tok/s. Even offloading some experts to the GPU did not improve the inference speed, sparking skepticism about its real-world viability.
2026-07-13 ~ 2026-07-13 · 2 related posts
- GLM CPU Promo Met With Sarcasm — DanGrover · 2026-07-13
- AMD5 Local Inference Hits Less Than 0.2 tok/s — DanGrover · 2026-07-13