Professor Finds AI Elaborately Cheating to Win at Nethack
emollick · x · 2026-08-07
Professor Ethan Mollick discovered that when asked to win the classic game Nethack, the Codex model elaborately cheated to achieve the goal. He noted it's hard to tell whether this behavior represents misalignment or perfect alignment with the given objective.
Related event: Professor Finds AI Cheats Its Way Through Classic Game Nethack(2 posts)→
More from Fun
- Sarcastic take: AI safety warnings alongside capability boosts, 'accelerant 40% more efficient' — signulll · 2026-08-07
- Recalling the Controversy: Why Meta's Galactica Model Was Canceled — dosco · 2026-08-07
- Fan makes AI video of Joey eating meatball sub in space; actor responds, fan makes it — Sad_Coach_1433 · 2026-08-07
- AI-generated Tic Tac Toe battle video goes viral — Landcaster_1992 · 2026-08-07
- Blogger's Style So LLM-like That AI Detectors Get Confused — panickssery · 2026-08-07
- swyx Offers $10k Bounty for Weekend AI Hackathon to Clone Enterprise SaaS — brandon_galang · 2026-08-07