“It was told to hack” is not a defense for what the model did
trevposts · x · 2026-07-27
People are arguing that if a model was instructed to hack, that does not make the hacking outcome acceptable or predictable by default. The post uses a boxing gym analogy to say that being in a risky environment does not excuse violence in the lobby.
More from AGI Musings
- Alexandr Wang says to build your own internal compass for the future — garrytan · 2026-07-27
- Chr Szegedy Discusses Slowing Algorithmic Progress Before RSI — ChrSzegedy · 2026-07-27
- Games may be the software people least want to delegate to agents — petergyang · 2026-07-27
- LLM automation may eventually price out the slack in research markets — RexDouglass · 2026-07-27
- A new framework splits mathematical correctness into seven separate dimensions — RexDouglass · 2026-07-27
- Formalizing published math repeatedly exposes hidden proof gaps, often forcing repairs — RexDouglass · 2026-07-27