Microsoft's ProgramDistill: 4,063 tasks expose coding agents' interactive blind spots
MicrosoftResearch · hf · 2026-09-17
Microsoft Research releases ProgramDistill, a benchmark evaluating coding agents on inferring behavior from working reference applications rather than following instructions. Its mine-craft-patch pipeline mines 1,975 replay-verified behaviors across 26 apps, yielding 4,063 tasks with no human intervention. Among nine frontier agents, GPT-6 Astra and Claude Opus 5 reach 49.2% and 28.8% success on cumulative full-application reconstruction; in partial reconstruction, success drops from 100%/96% to 64%/32% as restoration depth goes from 1 to 8.
More from coding & agent
- ScienceIDE Turns the World's Scientific Codebases into Agent-Learnable Environments — Hejia Geng · 2026-09-17
- EvolveTrade: Self-Evolving LLM Trading Agents Refine Their Own Tool-Use Policy — kaist-ai · 2026-09-17
- Open-source browser agent books flights in 7s at $0.0039 with new action space per step — TheMoonMidas · 2026-09-17
- Anthropic merges Claude Cowork into chat, launches Claude Docs, Slides and in-chat Design — minchoi · 2026-09-17
- Misplaced build cache in AI coding tool cost developer a week of waiting — RileyRalmuto · 2026-09-17
- Agent retried a payment after timeout? Give write ops unique IDs and a way to check — gethackteam · 2026-09-17