SWE-Touch shows coding agents fail when humans edit the code mid-task
CASIA · hf · 2026-08-04
- SWE-Touch benchmarks coding agents in a more realistic setting where users can inspect and modify code while the agent is still working.
- The benchmark introduces Counter-Edits: plausible user edits to task-relevant code that conflict with the agent’s goal.
- It mines task-critical regions from repair trajectories, generates user patches, and injects them with contextual messages when the agent reaches the relevant code.
- On SWE-bench Verified, the counter-edit setup lowers average resolve rate by 7.7 percentage points; the degradation also persists on SWE-Bench Pro and DeepSWE.
- The paper concludes that strong autonomous performance does not imply good workspace awareness, and highlights detecting edits, reconciling conflicts, and revalidating with tests as key future capabilities.
More from coding & agent
- SkillHone keeps full decision history to push agent scores up on GAIA by 15.8 — jiqizhixin · 2026-08-04
- OpenComputer launches an early preview of its serverless agents runtime — zeeg · 2026-08-04
- Fable hooks into GitHub Stacked PRs and launches a 27-PR workflow — dbreunig · 2026-08-04
- A meme says the orchestrator in agent workflows is becoming the human — vikvang1 · 2026-08-04
- A user is looking for MiniMax workflows that improve generation speed — PersonalMango2562 · 2026-08-04
- Which inference provider are you using in production, and what do you hate about it? — itsfabioroma · 2026-08-04