StateAct tops OSWorld 2.0 by using program state instead of screenshots
LiJunnan0409 · x · 2026-07-28
StateAct sets a new OSWorld 2.0 SOTA by relying less on pixels
A new paper, StateAct, reports a new best result on OSWorld 2.0 by flipping the usual computer-use recipe: instead of staring harder at screenshots, the agent treats program state as the primary interface and calls a GUI specialist only when a subgoal is truly visual.
Reported gains on the same Claude Opus 4.8
- Binary: 20.6% → 26.9%
- Partial: 54.8% → 61.6%
- Cost: about $72/task → $7.8/task
- Tokens: 224K → 100K
- GUI specialist is used for only 1.1% of steps
Broader benchmark lift
The same harness also improves Opus 4.8 across multiple benchmarks, including:
- OSWorld-Verified: 80.9 → 81.9
- WindowsAgentArena: 41.6 → 50.6
- AndroidWorld: 69.0 → 81.9
- MobileWorld: 51.3 → 70.1
The paper argues that computer-use is not just a vision problem: it is an agent reasoning problem that must connect what the system sees, the state it acts on, and the plan it maintains over long horizons.
Related event: Salesforce's StateAct Cuts Agent Costs 9x by Prioritizing State Over Pixels(4 posts)→
More from coding & agent
- ChatGPT Voice plus Codex lets users steer coding agents hands-free all day — nickbaumann_ · 2026-07-28
- MazeBench is a 3D benchmark where today’s best agents still fail the first levels — JasonBotterill · 2026-07-28
- CTI Expert turns Claude into a cyber threat intelligence analyst with 74+ commands — tom_doerr · 2026-07-28
- Agent Mini offers a 3,000-line local-first AI agent with shell, memory, and vision — Lordrovks · 2026-07-28
- Agensis pitches a shared workspace where AI agents reply by default and keep memory — jasonkneen · 2026-07-28
- Demo shows an Android app workflow using Antigravity, BigQuery, and Cloud Run — rseroter · 2026-07-28