GLM-5.2’s blog hints Z.ai dropped GRPO and went back to PPO
bycloud · youtube · 2026-07-24
GLM-5.2’s blog suggests Z.ai moved back from GRPO to PPO
The video argues that the most interesting part of GLM-5.2 is not just the benchmark performance, but the training details in Z.ai’s blog.
Key takeaway:
- the team reportedly decided to stop using GRPO
- and go back to PPO instead
The post frames this as a useful signal for how the team thinks about RL optimization, beyond the model’s headline results. The video is also sponsored and includes unrelated promo links, but the substantive point is the training-strategy shift discussed in the GLM-5.2 blog.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11