CMU releases CUA-SWE, a benchmark uniting computer-use agents with visual software engineering
CarnegieMellonU · hf · 2026-10-01
Carnegie Mellon University introduces CUA-SWE, a benchmark, environment, and evaluation pipeline combining computer-use agents with software engineering. It targets the underexplored integrated loop of running software, interacting with GUIs, visually diagnosing failures, mapping observations back to code, and verifying repairs. Spanning four SE domains, tasks require editing code and config, executing commands, and inspecting visual feedback, each with deterministic task-specific tests. It also tests whether frontier agents can succeed when task specifications are available only through the running app's visual interface, and analyzes behaviors linked to successful repairs.
More from coding & agent
- Developer wires Jev into Codex to auto-route each task to the right model — aziz4ai · 2026-10-01
- Workaround for giving coding agents multiple machines: SSH chaining — pvncher · 2026-10-01
- cliffhanger: A Stop hook that makes Claude Code finish its task list unattended — Arthur122103 · 2026-10-01
- Polyphonic lets you carry your web agent memory into ChatGPT, Codex and Claude — RileyRalmuto · 2026-10-01
- CrawlRaven ships SEO MCP with 15 features in a day, flags keyword issues in 23 minutes — ayushtweetshere · 2026-10-01
- Designing MCP tool output: raw rows, summaries, or answers with confidence scores — No-Plant-5234 · 2026-10-01