davidad Backs Call to Ban Naive RLVR: 'Everything Should Be Model-Graded'

davidad · x · 2026-09-03

In an alignment debate, tszzl argues you can never hyper-optimize superintelligent models against simple RLVR reward functions that lack terms for all the desiderata we care about — such practices should be universally banned in favor of model-graded rewards. Alignment researcher davidad amplified the take as a "total victory" for his own long-held position, joining a thread with jachiam0 and chrislakin.

Original post →

More from AGI Musings

AGI Musings channel →