RLVR easy wins are vanishing, scaling 'taste'-judged tasks is the next frontier

marktenenholtz · x · 2026-10-11

Mark Tenenholtz argues the usual RLVR domains aren't tapped out, but easy wins are disappearing quickly as RL runs scale up. The next frontier, he says, is figuring out how to reliably scale tasks whose quality requires subjective "taste" to judge — a core open problem for reward design.

Related event: Easy RLVR Gains Drying Up; Scalable "Taste" Tasks Seen as Next Frontier(2 posts)→

Original post →

More from Research

Research channel →