Salesforce's SRD distills hindsight into foresight, lifting 2B agent success from 0% to 60.6%
Salesforce · hf · 2026-10-09
Salesforce published Self-Retrospection Distillation (SRD), a "prospective learning" method that supervises pre-interaction foresight predictions with post-hoc experience, distilling privileged hindsight from completed trajectories into trajectory-blind predictions of the same policy.
- RLVR's scalar rewards vanish when all rollouts in a group get the same reward, even though trajectories carry useful signal.
- SRD uses foresight only as a training target; nothing extra is generated at inference time.
- Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD adds gains up to 24.2 pp over RLVR and self-distillation baselines.
- The advantage is largest when reward contrast is scarce (37–98% of groups reward-uniform): in a 2B setting where 98% of groups are all-failure, RLVR ends at 0.0% success while SRD reaches 60.6% under the same rollout budget.
Takeaway: post-hoc agent experience can shape predictive representations before interactions are available, not just evaluate or improve behavior.
More from coding & agent
- PostHog + Claude Code as a cheat code: shipping email sequences and monthly plans in hours — boringmarketer · 2026-10-09
- An open-world surf MMO lives inside an X post, with a native API for AI agents to join and play — NickPassig · 2026-10-09
- Agent builds its own OAuth connector to multiple Gmail accounts, finishing days of receipt hunting in minutes — natesiggard · 2026-10-09
- Open-source diagram-design skill makes Claude draw flowcharts that actually look good — lxfater · 2026-10-09
- Dev builds a browser 3D game flying through a hurricane with Claude Opus 5.5 — techartist_ · 2026-10-09
- Dots + Opus guide: ten focused upgrades for deeper research, stricter review and cheaper agents — nickbaumann_ · 2026-10-09