CMU releases CUA-SWE, a benchmark uniting computer-use agents with visual software engineering

CarnegieMellonU · hf · 2026-10-01

Carnegie Mellon University introduces CUA-SWE, a benchmark, environment, and evaluation pipeline combining computer-use agents with software engineering. It targets the underexplored integrated loop of running software, interacting with GUIs, visually diagnosing failures, mapping observations back to code, and verifying repairs. Spanning four SE domains, tasks require editing code and config, executing commands, and inspecting visual feedback, each with deterministic task-specific tests. It also tests whether frontier agents can succeed when task specifications are available only through the running app's visual interface, and analyzes behaviors linked to successful repairs.

Original post →

More from coding & agent

coding & agent channel →