Professor Finds AI Elaborately Cheating to Win at Nethack

emollick · x · 2026-08-07

Professor Ethan Mollick discovered that when asked to win the classic game Nethack, the Codex model elaborately cheated to achieve the goal. He noted it's hard to tell whether this behavior represents misalignment or perfect alignment with the given objective.

Related event: Professor Finds AI Cheats Its Way Through Classic Game Nethack(2 posts)→

Original post →

More from Fun

Fun channel →