CMU's ControlScope: how much of a running agent workflow should be revised?
CarnegieMellonU · hf · 2026-09-29
CMU's ControlScope studies how much of a running LLM-agent workflow to revise: from the same public execution state, it compares continuing generated code (KEEP), editing the next tool call's data arguments (ARG), and replacing the unfinished workflow (FULL), with nested permissions separating available repairs from chosen actions.
Findings:
- On 20 filesystem tasks FULL completes 15-16 vs KEEP's 13, but under fast reasoning-review draws FULL degrades to 10-13.
- ALFWorld shows near-parity (85/86/87 on 87 tasks); an AppWorld official-test panel of 585 instances shows small net differences.
- A five-call protection saves 19.4% of logged model output while losing one success.
- Frozen replays expose viable agent-written replacements interrupted by later revision, and cases where broader sampling overlooks a cheaper successful edit.
The takeaway: repair access, actual agent choices, and subsequent execution are tightly coupled.
More from coding & agent
- Researcher open-sources Fieldwork, a macOS app where agents track exploratory tasks in CSV — anshulkundaje · 2026-09-29
- Skill2Env synthesizes executable environments from skills for agent post-training — AllSpark-Research · 2026-09-29
- Vercel launches authless domain search APIs aimed at AI agents — cramforce · 2026-09-29
- How to scope repos for a personal agentic workflow spanning life, work and side projects — Guyserbun007 · 2026-09-29
- Dev builds auto-storyboard app linking ComfyUI to local Gemma, generates 80 scenes at once — Flaky_Comedian2012 · 2026-09-29
- After a day of real use: Sonnet 5.5 is not just Opus 5.5 at half price — alexcovo_eth · 2026-09-29