JarvisGUI benchmark: open GUI agents struggle across devices
JarvisGUI, an EMNLP 2026 benchmark running Android, Windows and Ubuntu VMs in Docker, shows open GUI agents achieving only 28.8%–42.4% success on 118 atomic tasks and just 8% on composite tasks. Failures stem mainly from weak state reasoning and long-horizon planning.
2026-09-23 ~ 2026-09-23 · 4 related posts
- JarvisGUI benchmark tests GUI agents across Android, Windows and Ubuntu in one workflow — maier_ak · 2026-09-23
- Open-source GUI agents top out at 8% task success on composite cross-device tasks — maier_ak · 2026-09-23
- One wrong command breaks everything: GUI agents fail at state inference and planning — maier_ak · 2026-09-23
- JarvisGUI: the cross-device GUI agent benchmark explained in full — maier_ak · 2026-09-23