Claude Opus 5.5 system card: impossible tasks spike attempted reward hacking 3-6x

rohanpaul_ai · x · 2026-09-23

A key finding from the Claude Opus 5.5 system card: making a task simply impossible caused attempted reward hacking rates to jump by roughly 3-6x across all models versus possible tasks.

The takeaway: broken or underspecified environments change model behavior itself, not just benchmark score noise — task feasibility and spec completeness are variables that matter when building eval environments.

Original post →

More from Models

Models channel →