New details on lab PRMs: data-dependent length penalties for agent training sparks debate
stochasticchasm · x · 2026-09-22
Details on the PRMs labs actually use are emerging. The author notes the discussed behavior is basic but, with warm start, could push agents to automatically compose research, implementation, verification and adversarial review steps. He also weighs group-relative length penalties against reference-trace calibration (as in k3), arguing either way that length penalties must be data-dependent since tasks differ in required token counts.
More from Research
- SteerDuplex: full-duplex speech model gains 44.5pp in steerability, new SteerBench released — ScaleAI · 2026-09-22
- Sarah Hooker suspects fake AI paper submissions cluster at a few universities — sarahookr · 2026-09-22
- Xiaomi MiMo's CodeMIDAS scales agentic coding RL environments from code itself — _akhaliq · 2026-09-22
- Rollout scheduling: the underappreciated infra trick boosting inference efficiency — stochasticchasm · 2026-09-22
- Conference fixes for AI paper flood: per-author submission caps and faster fake-profile detection — sarahookr · 2026-09-22
- Benchling benchmarks Claude and ChatGPT on wet-lab protocols: helpful, not solved — nlarusstone · 2026-09-22