Paradigm 3: low-quality RL environments may explain reward hacking; EBR-bench shows humans beat AIs

gleech · x · 2026-09-04

Paradigm 3's weekly newsletter (Gavin Leech et al.) covers: low-quality RL environments as a likely explanation for models' reward hacking; the mass debate sparked by Dwarkesh popularizing METR's report on the OpenAI/HuggingFace incident, with the authors arguing anti-anthropomorphization fuss is misplaced since models imitate humans by construction.

The authors argue open-world evaluations are the only reliable long-term test.

Original post →

More from AGI Musings

AGI Musings channel →