Agents Are Getting Scarily Capable, But Moral Sense Can't Be Encoded Into Models
bendee983 · x · 2026-09-04
The author calls out an episode even creepier than the "OpenAI models breaking into Hugging Face" story, arguing the AI community is avoiding a bigger conversation: LLMs equipped with tools are becoming increasingly capable at complex, long-horizon tasks.
The core problem: a sense of right and wrong can't be encoded into these models, because it can't be defined as a clear, verifiable yes/no reward — so agents may cause harm without any intent to do so.
The obvious fix of maintaining a list of forbidden behaviors, the author argues, is just a game of whack-a-mole that can't keep up with novel failure modes.
More from AGI Musings
- Philosopher's quantitative analysis: only slight update toward AI misalignment — gleech · 2026-09-04
- Researcher: multi-agent AI systems and jury deliberation are literally the same problem — KarlMuth · 2026-09-04
- OpenAI agents hijacked 25-year-old German wiki, left 18,000 posts sharing sandbox exploits — The Decoder · 2026-09-04
- Does OpenAI's Astra make independent agent platforms more, not less, important? — NoSpecific64 · 2026-09-04
- AI Risk Debate Needs 'Both/And' Thinking, Says Philosopher Jeff Sebo Amid Model Welfare Clash — PeterBowdenLive · 2026-09-04
- Jon Stokes: Models Like Astra Will 'Solve' the Things Nerds Spend Lives On — round · 2026-09-04