Researchers lament: no good RL environment sets exist that allow diffing behaviors across training runs

1a3orn · x · 2026-09-23

In a technical exchange, @menhguin argues such information is unusable without an immediate before/after diff, since current methods can't tell which features belong to which training phase. @1a3orn agrees a diff is needed, saying the work likely waits on a good, minimal but non-stupid set of RL environments that produce both good behavior and reward-hacking behavior — and such a set simply doesn't exist yet.

Original post →

More from Research

Research channel →