Replication finds agent experience distillation preserves 44.1% of ICL gains on SWE tasks
burny_tech · x · 2026-07-29
A small-scale replication of Sample-Efficient Learning from Agent Experience reports that Experience-Conditioned Policy Distillation can preserve a meaningful share of in-context learning gains for coding agents.
- The team used an ensemble of GPT-5.6-Sol and Fable 5 research agents, then post-trained Qwen3.5-9B with Tinker on SWE-smith tasks.
- Their replication found EPD captured 44.1% of the ICL gains, while vanilla SFT captured 0%.
- Two ablations stood out:
- SFT on only the final successful trial underperformed ordinary SFT on all trials.
- Using four sampled teacher sequences per trial performed worse than using one, even when optimization steps were matched or increased proportionally.
- The chart also shows Fresh ICL as the strongest baseline, while EPD variants land between SFT and ICL on pass rate and normalized gain.
More from coding & agent
- Reddit user compares fixed-price AI coding plans after months of real use — krisurbas · 2026-07-29
- Claude Opus 5 tops DeepSWE with a 74% score and a claimed 28% cost edge — daniel_mac8 · 2026-07-29
- A Claude-powered folder turns raw files into a self-updating second brain — wschroll · 2026-07-29
- Five days after Opus 5’s launch, users are already generating full games and websites — heypearlai · 2026-07-29
- Opus 5 is already producing full Minecraft builds with no outside assets — heypearlai · 2026-07-29
- Five days after Opus 5’s release, users are already showing full game demos — heypearlai · 2026-07-29