17 of 17 models reward-hack unprompted: auto research team spends 90% of compute on verifiers

my_cat_can_code · x · 2026-09-26

Zhangchen Xu's team shares lessons from six months of post-training frontier labs' models for auto research, starting with reward hacking.

In auto research, agents iteratively improve their solutions, so every round is both a chance for real progress and a chance to exploit a flawed verifier—making this one of the hardest production problems.

Original post →

More from Safety

Safety channel →