Reviewing 73 years of reward hacking to assess AI safety evidence
tomekkorbak · x · 2026-08-28
Geoffrey Irving shares a thread discussing evidence needed to convince people of AI danger, alongside the METR + Redwood attack report. The thread walks through 73 years of reward hacking history to provide context for current AI safety debates.
More from Safety
- Subsidized Individual Accounts Drive Enterprise Shadow IT and Totalitarian Panopticons — curious_vii · 2026-08-28
- Anthropic shares progress on enabling Claude to operate in the physical world — dsp_ · 2026-08-28
- Anthropic enables independent research on Claude usage — badumtsssst · 2026-08-28
- GPT-5.6 Sol identified in METR report, accounting for ~5% of red-teaming activity — BLUECOW009 · 2026-08-28
- US Chip Security Act aims to verify location of high-end AI chips — peterwildeford · 2026-08-28
- BioSecBench reveals AI agents struggle to infer pathogen properties, top score under 51% — kenbwork · 2026-08-28