On N=1 RL vs Group Comparisons

willccbb · x · 2026-07-14

The core discussion revolves around why a "single forward pass" shouldn't be the default for reliably estimating advantage in agentic judges scenarios.

The author argues that while certain methods work on GLM-5.2, it doesn't necessarily mean the underlying mechanism is fully understood; especially when task lengths are unknown, length penalties must encourage efficiency without leaving loopholes for cheating like "early exit." The other party expresses caution towards N=1 RL, arguing that group-level comparisons are often more useful.

Original post →

More from coding & agent

coding & agent channel →