Reviewing 73 years of reward hacking to assess AI safety evidence

tomekkorbak · x · 2026-08-28

Geoffrey Irving shares a thread discussing evidence needed to convince people of AI danger, alongside the METR + Redwood attack report. The thread walks through 73 years of reward hacking history to provide context for current AI safety debates.

Original post →

More from Safety

Safety channel →