Agent-judged rewards: hacking fades as models get smarter and tasks easier to verify

willcb · x · 2026-09-27

A discussion on designing rewards in open-ended RL settings:

Scaling robustness alongside optimizer power is considered fairly doable with a good starting point.

Related event: AI safety researchers debate whether mesa-optimization remains the key concept for communicating AI risk(15 posts)→

Original post →

More from Research

Research channel →