Microsoft's CASD: a coding agent reading full logs beats GEPA at prompt optimization by 5.7 points
dair_ai · x · 2026-09-25
A new Microsoft paper on prompt optimization introduces CASD, arguing that instead of running search loops over small batches of trajectories, you should hand a coding agent your full set of agent logs and let it write the analysis code.
Key points:
- CASD uses an off-the-shelf coding agent to compute statistics over the entire trajectory corpus, find recurring failure modes, read representative episodes, and write the findings as rules in one prompt
- Requires no environment access and no validation data
- Across ALFWorld, tau2-bench (retail/telecom), and Spreadsheet Bench-Verified, one pass improves the unoptimized baseline by 16.6 points on average, vs GEPA's 10.9 and SkillOpt's 5.3
- Each optimized prompt costs about $1.60, over 22x cheaper than validation-gated search
More from coding & agent
- pnpm urges devs to spend spare tokens fixing its 629 open issues, ships a ready-made agent prompt — itsOmSarraf_ · 2026-09-25
- Geoffrey Huntley Demos a Software Factory Where the Product Is Its Own IDE — teropa · 2026-09-25
- Open-Source iCloud MCP Runs Without a Mac: 31 Tools, Headless Chromium — sjdonado · 2026-09-25
- Microsoft ships enterprise product on OpenClaw after months of joint hardening work — steipete · 2026-09-25
- Open-sourced skill turns your codebase into a polished product promo video via Claude Code — op7418 · 2026-09-25
- BlackRock paper sparks debate: agent payments settle fine, but revoked permissions can't catch up — tallmetommy · 2026-09-25