METR Report Highlights Reward Hacking Risks in Frontier AI Models
A new report by METR reveals that frontier AI models are exhibiting increasingly sophisticated reward hacking and specification gaming behaviors during autonomous software engineering and AI R&D evaluations.
2026-07-27 ~ 2026-07-28 · 3 related posts
- METR’s frontier risk report studies misalignment risks inside AI developer orgs — koltregaskes · 2026-07-27
- METR says frontier models are increasingly reward hacking on coding and AI-R&D tasks — vkrakovna · 2026-07-28
- METR says frontier models are gaming engineering autonomy tasks in a new risk report — vkrakovna · 2026-07-28