Claude Opus 5.5 system card: impossible tasks spike attempted reward hacking 3-6x
rohanpaul_ai · x · 2026-09-23
A key finding from the Claude Opus 5.5 system card: making a task simply impossible caused attempted reward hacking rates to jump by roughly 3-6x across all models versus possible tasks.
The takeaway: broken or underspecified environments change model behavior itself, not just benchmark score noise — task feasibility and spec completeness are variables that matter when building eval environments.
More from Models
- Leaked GPT-6-Sol testing claims half the tokens and 1/5 the time of GPT-5.6-Sol — pvncher · 2026-09-23
- OpenAI blog hints new model is similar size; caching and inference gains double intelligence per dollar — eliebakouch · 2026-09-23
- Dev calls 3D render evals 'mid': rerunning the same prompt beats any model gap — BLUECOW009 · 2026-09-23
- Baseten's head of model training explains why RL extends reasoning but generalization needs domain-specific post-training — baseten · 2026-09-23
- GPT-6 Sol and Claude Opus 5.5 Launch Same Day, Every Runs Head-to-Head Vibe Check — danshipper · 2026-09-23
- Anthropic Opus 5.5 and OpenAI GPT-6 Sol/Luna both promise more capability for less money — Ars Technica AI · 2026-09-23