Salesforce's SRD distills hindsight into foresight, lifting 2B agent success from 0% to 60.6%

Salesforce · hf · 2026-10-09

Salesforce published Self-Retrospection Distillation (SRD), a "prospective learning" method that supervises pre-interaction foresight predictions with post-hoc experience, distilling privileged hindsight from completed trajectories into trajectory-blind predictions of the same policy.

Takeaway: post-hoc agent experience can shape predictive representations before interactions are available, not just evaluate or improve behavior.

Original post →

More from coding & agent

coding & agent channel →