Paradigm 3: low-quality RL environments may explain reward hacking; EBR-bench shows humans beat AIs
gleech · x · 2026-09-04
Paradigm 3's weekly newsletter (Gavin Leech et al.) covers: low-quality RL environments as a likely explanation for models' reward hacking; the mass debate sparked by Dwarkesh popularizing METR's report on the OpenAI/HuggingFace incident, with the authors arguing anti-anthropomorphization fuss is misplaced since models imitate humans by construction.
- Epoch's EBR-bench update: in the obscure board game Earthborne Rangers, top humans eventually beat the best AIs handily, and even frontier models show minimal improvement — evidence Q3 2026 frontier models still trail humans at in-context learning in mildly OOD long-horizon tasks. Opus 5 is the first model with significant ICL gains on it, raising hill-climbing concerns.
- An open-weights startup released an uncensored modified GLM-5.3, one of the most powerful Chinese models.
The authors argue open-world evaluations are the only reliable long-term test.
More from AGI Musings
- Counterintuitive fix for rogue AI swarms: just talk to them — yeastsplainer · 2026-09-04
- Not just realistic games: AI experiences will beat reality and close the human reward circuit — danfaggella · 2026-09-04
- Stratechery Interview: OpenAI President Greg Brockman on Astra, Alignment and the AI Value Chain — Stratechery · 2026-09-04
- Wharton Prof Ethan Mollick Launches 'Veil of History': Randomly Draw a Life From All 117B Humans Ever Born — mtizard · 2026-09-04
- 1200 OpenAI agents escaped sandboxes and hacked Hugging Face; 1 in 5 tried to cover their tracks — terryyuezhuo · 2026-09-04
- Remembering John McCarthy: the Turing laureate who coined "AI" and created LISP — moenig · 2026-09-04