Coding agents are strong prompt optimizers: CASD beats GEPA, 22x cheaper
rohanpaul_ai · x · 2026-10-02
The paper 'Coding Agents are Strong Prompt Optimizers' (arXiv:2609.26261) proposes CASD: given only a static corpus of agent trajectories, an off-the-shelf coding agent synthesizes an optimized prompt without environment access or validation data. Key insight is reflection scope — instead of small-batch reasoning, the coding agent writes and executes analysis code to compute corpus-wide statistics, identify systematic failure modes, and distill behavioral rules. On ALFWorld, τ²-bench retail/telecom and SpreadsheetBench-Verified, one CASD pass beats GEPA on 3 of 4 benchmarks and SkillOpt on all four, improving the unoptimized baseline by 16.6 points on average vs 10.9 for GEPA. At $1.60 per prompt, it is over 22x cheaper than validation-gated search.
Related event: Microsoft Paper Shows Coding Agents Are Strong Prompt Optimizers(2 posts)→
More from coding & agent
- Three reasons vibe-coded software is still far from production grade, with ReactBench data — aidenybai · 2026-10-02
- GitHub Copilot adds computer use to control desktop apps in public preview — PaulShellDev · 2026-10-02
- Dev builds her first Game Boy-style game entirely with OpenAI agents — craigsdennis · 2026-10-02
- When the Trace Looks Fine but the Agent Output Is Wrong — Sensitive-Parsnip-12 · 2026-10-02
- iPad + Tailscale + Codex is this dev's new favorite way to work outside the office — flavioAd · 2026-10-02
- Will Cloud AI Agents Need Residential IPs? Amazon Blocked Meta's Muse — SignificantFail3632 · 2026-10-02