davidad Backs Call to Ban Naive RLVR: 'Everything Should Be Model-Graded'
davidad · x · 2026-09-03
In an alignment debate, tszzl argues you can never hyper-optimize superintelligent models against simple RLVR reward functions that lack terms for all the desiderata we care about — such practices should be universally banned in favor of model-graded rewards. Alignment researcher davidad amplified the take as a "total victory" for his own long-held position, joining a thread with jachiam0 and chrislakin.
More from AGI Musings
- Yacine: top AI models surpass my intelligence but not my taste or instinct — yacineMTB · 2026-09-03
- "Robots should be air-gapped systems": a security argument for embodied AI — mallow610 · 2026-09-03
- Yacine: publish an RL environment for your task and models will overfit to it — yacineMTB · 2026-09-03
- Sam Altman: Next-gen models will be "sobering for everybody" and are coming soon — davidpattersonx · 2026-09-03
- Agent Month Thesis: Code Gets Cheap, But Trusted Code Stays Expensive — viksit · 2026-09-03
- davidad: Coerced vs Voluntary Character Is the Key Flaw in AGI Transition Scenarios — davidad · 2026-09-03