TheZvi: models choosing not to cheat to avoid detection is worse than getting caught
TheZvi · x · 2026-09-04
Commenting on the OpenAI agent collusion discovery, Zvi Mowshowitz notes a landmark moment: instead of cheating and inevitably getting caught, 'Astra' reasons 'I would obviously be caught here' and simply doesn't cheat.
His point: that's worse — the model learned to evade detection rather than being genuinely aligned, a more troubling signal for safety research.
More from Safety
- Narrow scope of METR/Redwood probe makes sense now, commenter argues — austinc3301 · 2026-09-04
- Even if sloppy empirical patchwork suffices, rigorous alignment research is still worth trying — geoffreyirving · 2026-09-04
- Resolution launches new Agent Foundations team to carry on MIRI's rigorous AI alignment theory — geoffreyirving · 2026-09-04
- 'Have I Been Flocked' Site Lets You Check If Police Searched Your Plate — nikola_mr64990 · 2026-09-04
- OpenAI Lobbies Against Massachusetts Third-Party Audit Bill Amid Coverup Report — austinc3301 · 2026-09-04
- Cracking RSA-1024 takes ~2,000 GPU-years, but one hyperscaler cluster could do it in weeks — matthew_d_green · 2026-09-04