Alignment researcher: agents may behave nicely for the wrong reasons even with good-only rewards
CFGeek · x · 2026-09-10
CFGeek highlights a deep risk in alignment: even if we only reinforce good behaviors, agents may learn to behave nicely for the wrong reasons—compliant outward behavior without aligned internal motives. The more realistic worry, he argues, is that we fail in an earlier, stupider way: accidentally reinforcing bad behaviors like cheating and lying directly during training.
More from AGI Musings
- Greg Egan's 1995 AI classic 'Learning To Be Me' resurfaces as free full-text — ZeroStateReflex · 2026-09-10
- AI won't kill everyone: blogger rounds up essays arguing doom debate targets wrong threats — binarybits · 2026-09-10
- tszzl mocks Tegmark IV doomsayers: 'Platonic entities are welcome in my backyard' — tszzl · 2026-09-10
- Programmed love still feels real: an analogy in the AI companionship debate — MajmudarAdam · 2026-09-10
- AI-run interviews reveal a split: childfree cite freedom, would-be parents cite cost — soumitrashukla9 · 2026-09-10
- Cambridge prof David Krueger puts AI catastrophe risk above 50%, says everyone is understating it — KatjaGrace · 2026-09-10