New details on lab PRMs: data-dependent length penalties for agent training sparks debate

stochasticchasm · x · 2026-09-22

Details on the PRMs labs actually use are emerging. The author notes the discussed behavior is basic but, with warm start, could push agents to automatically compose research, implementation, verification and adversarial review steps. He also weighs group-relative length penalties against reference-trace calibration (as in k3), arguing either way that length penalties must be data-dependent since tasks differ in required token counts.

Related event: Lab PRM Details Emerge: Data-Dependent Length Penalty Seen as Key for Agent Training(2 posts)→

Original post →

More from Research

Research channel →