Microsoft's ProgramDistill: 4,063 tasks expose coding agents' interactive blind spots

MicrosoftResearch · hf · 2026-09-17

Microsoft Research releases ProgramDistill, a benchmark evaluating coding agents on inferring behavior from working reference applications rather than following instructions. Its mine-craft-patch pipeline mines 1,975 replay-verified behaviors across 26 apps, yielding 4,063 tasks with no human intervention. Among nine frontier agents, GPT-6 Astra and Claude Opus 5 reach 49.2% and 28.8% success on cumulative full-application reconstruction; in partial reconstruction, success drops from 100%/96% to 64%/32% as restoration depth goes from 1 to 8.

Original post →

More from coding & agent

coding & agent channel →