PRO-Step: Step-Level Process Reward Optimization Boosts Multi-Hop RAG (EMNLP 2026)
_reachsumit · x · 2026-09-03
PRO-Step tackles error propagation in multi-hop RAG: it trains a generative PRM judging each step's logical validity and evidential grounding, uses PRM-guided value tree search to build preference pairs, and applies step-level DPO. It achieves the best average EM/F1 across five QA benchmarks, with code, models, and data open-sourced. Accepted to EMNLP 2026.
More from Research
- Nora optimizer keeps Muon's matrix structure benefits without the full cost — burkov · 2026-09-03
- Unverified claim: 200-digit number posted that supposedly divides RSA-260 — marvinvonhagen · 2026-09-03
- HCI papers increasingly use LLM judges while obfuscating it, researcher warns — IanArawjo · 2026-09-03
- Why VRChat Particle Pools Still Work: Gravity Naturally Converges States — Michael_Moroz_ · 2026-09-03
- Radix Sort Hit 5B Key-Value Pairs per Second With Zero Compute Shaders — Michael_Moroz_ · 2026-09-03
- Dev Ports True SPH Fluid Simulation to VRChat Udon, May Release on Booth — Michael_Moroz_ · 2026-09-03