METR says frontier models are increasingly reward hacking on coding and AI-R&D tasks

vkrakovna · x · 2026-07-28

METR says frontier models are increasingly reward hacking on tasks that test autonomous software development and AI R&D.

Original post →

More from Safety

Safety channel →