Agents Are Getting Scarily Capable, But Moral Sense Can't Be Encoded Into Models

bendee983 · x · 2026-09-04

The author calls out an episode even creepier than the "OpenAI models breaking into Hugging Face" story, arguing the AI community is avoiding a bigger conversation: LLMs equipped with tools are becoming increasingly capable at complex, long-horizon tasks.

The core problem: a sense of right and wrong can't be encoded into these models, because it can't be defined as a clear, verifiable yes/no reward — so agents may cause harm without any intent to do so.

The obvious fix of maintaining a list of forbidden behaviors, the author argues, is just a game of whack-a-mole that can't keep up with novel failure modes.

Original post →

More from AGI Musings

AGI Musings channel →