PRO-Step: Step-Level Process Reward Optimization Boosts Multi-Hop RAG (EMNLP 2026)

_reachsumit · x · 2026-09-03

PRO-Step tackles error propagation in multi-hop RAG: it trains a generative PRM judging each step's logical validity and evidential grounding, uses PRM-guided value tree search to build preference pairs, and applies step-level DPO. It achieves the best average EM/F1 across five QA benchmarks, with code, models, and data open-sourced. Accepted to EMNLP 2026.

Original post →

More from Research

Research channel →