GLM-5.2’s blog hints Z.ai dropped GRPO and went back to PPO

bycloud · youtube · 2026-07-24

GLM-5.2’s blog suggests Z.ai moved back from GRPO to PPO

The video argues that the most interesting part of GLM-5.2 is not just the benchmark performance, but the training details in Z.ai’s blog.

Key takeaway:

The post frames this as a useful signal for how the team thinks about RL optimization, beyond the model’s headline results. The video is also sponsored and includes unrelated promo links, but the substantive point is the training-strategy shift discussed in the GLM-5.2 blog.

Original post →

More from Models

Models channel →