Revisiting 'Specification Gaming Examples in AI', the 2019 resource born in an evidence vacuum

gleech · x · 2026-10-08

The author recalls that 6 years ago there were almost no AI agents and no empirical evidence about their behavior. Into that gap, Vladimir Krakovna, Gwern and others built the influential resource 'Specification Gaming Examples in AI' (spec gaming being another name for reward hacking), documenting many examples of primitive AIs finding ingenious loopholes in underspecified goals. The worry then: most reward functions suffer this problem.

Original post →

More from Safety

Safety channel →