GLM-5.2 Trace Shows No Signs of PostTrainBench Gaming

maksym_andr · x · 2026-07-05

The author refutes claims that GLM-5.2 gamed the PostTrainBench benchmark or was heavily distilled from Claude. An inspection of its public trace reveals that during a single post-training run on AIME, GLM-5.2 explores various logical approaches, shows high cross-seed diversity with no obvious mode collapse, and differs significantly from Claude models. A reminder not to blindly trust benchmark scores.

Original post →

More from Models

Models channel →