Dense Process Supervision for Search Agents via Fact Utility Estimation

_reachsumit · x · 2026-09-02

The paper 'Dense Process Supervision for Search Agents via Fact Utility Estimation' proposes a method to enable dense process supervision for RL-based search agents. It addresses the credit assignment problem in outcome-reward RL by modeling reasoning as the accumulation of discrete evidence facts. The method extracts structured facts into a fact store, clusters semantically equivalent facts, and infers their posterior utility via Bayesian estimation. These utilities are converted into dense step-level rewards for RL training. Experiments on seven QA benchmarks show consistent outperformance over baselines.

Original post →

More from coding & agent

coding & agent channel →