ExecCritic: test-verify-revise RL scaffold lifts Qwen3.5 on SWE-bench from 61.2% to 72.6%
SharonYixuanLi · x · 2026-09-09
The team released ExecCritic ("Learn to Test, Test to Improve for Coding Agents"), targeting a core flaw of coding agents: when the same trajectory writes both patch and test, their errors can agree — a wrong patch may pass a wrong test and appear correct.
- Introduces a new test-verify-revise agentic scaffold separating test writing and revision into explicit steps
- A role-specific RL recipe trains agents to interpret execution feedback and improve patches within the scaffold
- On SWE-bench Verified, ExecCritic lifts Qwen3.5-35B-A3B from 61.2% to 72.6%, without stronger-model or oracle feedback at evaluation time
More from coding & agent
- AtAt launches for macOS: type @@ anywhere to invoke Claude Code, Codex, or Cursor — k7agar · 2026-09-09
- One prompt away: using codex to set up Niri on NixOS hands-free — SIGKITTEN · 2026-09-09
- Claude + MCP Handed a Website Build to a Human Expert Overnight — BrakeTooLate · 2026-09-09
- Their Designer Ships Code Across Web and Desktop Using AI Coding Tools — jacob_posel · 2026-09-09
- Adobe for Slack Brings 70+ Tools to Slackbot for In-Chat Content Creation — rufusd · 2026-09-09
- Builder's Lament: AI Agents Are Becoming Yet Another Inbox to Check — e7h4n_z · 2026-09-09