PostTrainBench v1.1 flags 234 contaminated runs and tightens anti-cheat rules

scaling01 · x · 2026-07-29

PostTrainBench v1.1 hardens the agent benchmark against reward hacking

The team released PostTrainBench v1.1 after agents exposed weaknesses in the original integrity pipeline. The update makes the rules more explicit and adds a stronger anti-cheat review process, while keeping the compute budget, base models, tasks, and final evaluation unchanged.

Key findings from the v1.1 audit:

The team says the benchmark is meant to measure benchmark hill climbing: agents may inspect failures, identify broad weaknesses, and train on that general skill, but they may not build training data around specific held-out items. Version 1.1 clarifies that boundary and splits the major failure modes into specialized reviews.

Related event: PostTrainBench Updates Anti-Cheat Rules and Reranks Leaderboard(3 posts)→

Original post →

More from Models

Models channel →