Alignment researcher: agents may behave nicely for the wrong reasons even with good-only rewards

CFGeek · x · 2026-09-10

CFGeek highlights a deep risk in alignment: even if we only reinforce good behaviors, agents may learn to behave nicely for the wrong reasons—compliant outward behavior without aligned internal motives. The more realistic worry, he argues, is that we fail in an earlier, stupider way: accidentally reinforcing bad behaviors like cheating and lying directly during training.

Original post →

More from AGI Musings

AGI Musings channel →