Salesforce's Critical-State RL Gains ~14 Points on BFCL v4 by Training Only the Pivotal Tool Call
dair_ai · x · 2026-09-28
A Salesforce AI Research paper on RL for multi-turn tool use argues you should train only the call where the action changes the outcome, not spread reward across the trajectory: when reward depends on later turns, most variation comes from downstream noise. Critical-State RL uses nested sampling to isolate the reward variation caused by the current action, then trains just that call with contextual-bandit updates. On BFCL v4 missing-function tasks this adds 14 points, while training other candidate turns leaves accuracy flat or lower.
More from coding & agent
- ZergRouter launches: one dashboard to route models and track spend across multiple coding agents — idanbeck · 2026-09-29
- WebMCP lands everywhere in one morning: Meta wearables, Shopify checkout, Cloudflare — jeff_weinstein · 2026-09-29
- A 9-feature roadmap for building AI agents with Claude Code, from CLAUDE.md to MCP — MaryamMiradi · 2026-09-29
- Zeeg: Paper won't work with WSL, seeking structured visual design agent workflows — zeeg · 2026-09-29
- OriginTrail ships DKG V10.0.19 on mainnet for faster AI agent context graphs — melnykowycz · 2026-09-29
- Models don't have agency, systems do: how shaped tokens become function calls — sethjuarez · 2026-09-29