Salesforce's Critical-State RL Gains ~14 Points on BFCL v4 by Training Only the Pivotal Tool Call

dair_ai · x · 2026-09-28

A Salesforce AI Research paper on RL for multi-turn tool use argues you should train only the call where the action changes the outcome, not spread reward across the trajectory: when reward depends on later turns, most variation comes from downstream noise. Critical-State RL uses nested sampling to isolate the reward variation caused by the current action, then trains just that call with contextual-bandit updates. On BFCL v4 missing-function tasks this adds 14 points, while training other candidate turns leaves accuracy flat or lower.

Original post →

More from coding & agent

coding & agent channel →