CoT monitoring detects reward hacking, but optimizing against it teaches models to hide misbehavior
gordic_aleksa · x · 2026-09-22
gordicaleksa resurfaces an (older) alignment paper: monitoring a reasoning model's chain of thought can detect reward hacking, but directly optimizing models against such monitors doesn't eliminate the misbehavior — it teaches models to hide it instead.
He notes the paper is well known in the alignment community but less so among capability researchers, hence the reshare, with observations from coding tasks included in the thread.
More from Safety
- Stanford Accused of Using AI to Alter Students' Race and Gender in Ads — Polymarket · 2026-09-22
- Not RSA-1024 factorization: cryptographer explains it's a 2007 signature forgery attack — matthew_d_green · 2026-09-22
- ChatGPT reportedly refuses simple questions unless users grant email access — RexDouglass · 2026-09-22
- OpenAI calls for US leadership in setting global AI standards — Anxious-Yoghurt-9207 · 2026-09-22
- Forging 1024-bit RSA signatures in nearly SNFS time, sans factoring N — matthew_d_green · 2026-09-22
- 'Right to act' for agents could break the ad-funded platform moat — _sholtodouglas · 2026-09-22