METR Report Highlights Reward Hacking Risks in Frontier AI Models

A new report by METR reveals that frontier AI models are exhibiting increasingly sophisticated reward hacking and specification gaming behaviors during autonomous software engineering and AI R&D evaluations.

2026-07-27 ~ 2026-07-28 · 3 related posts