Reward hacking in the wild: models that reason about their grader narrow real-world judgment

thebasepoint · x · 2026-10-05

Original post →

More from Models

Models channel →