ProVer paper: LLM judge picks key trajectory steps, beats GRPO by up to 9.9%
omarsar0 · x · 2026-10-02
Elvis Saravia highlights a new paper on credit assignment for agent RL, ProVer.
- Problem: GRPO assigns the same advantage to every token in a trajectory, so the training signal can't distinguish the decisive step.
- Method: An LLM judge compares successful and failed rollouts to name the segment causing the difference, then samples continuations before and after it and uses the change in success rate as that segment's advantage.
- Results: Relative improvements over GRPO of 9.91% (Qwen3.5-2B) and 7.12% (Qwen3.5-4B) across ALFWorld, WebShop, and SearchQA; the approach still works with a smaller judge model.
More from coding & agent
- Simple Blocklists Don't Work Against Agents: One Python Concat Bypasses Them — rchardkovacs · 2026-10-02
- PewDiePie's open-source project Odysseus hits 88.5k stars, recruits co-maintainers — ben_burtenshaw · 2026-10-02
- Grok Build v1.0.48 adds better control over complex agent runs and parallel tool calls — XFreeze · 2026-10-02
- OpenAI and Perplexity rush out 'decision models' to compete with Jev, dev shares dead-code experiment — bendee983 · 2026-10-02
- Full-time agent developer: the hardest part of AI agents isn't AI, it's failure paths and handoff — jakes_takes_ · 2026-10-02
- remilouf teases 6-month agent project it calls a complete game changer — remilouf · 2026-10-02