MobileWorld benchmark: 201-task eval of phone GUI agents, best combo only ~52%
East-Muffin-6472 · reddit · 2026-09-07
A lit review of MobileWorld, a benchmark for autonomous mobile GUI agents that adds two novel axes: user-interaction tasks (agent must ask the GPT-4-simulated user for missing info) and MCP-augmented tasks (one-shot data retrieval via GitHub/arXiv tools instead of slow GUI taps).
- 201 tasks across 20 everyday apps, under 5% system apps (mostly open-source clones)
- Planner-executor setup: a VLM planner sees screenshots only and emits natural-language actions, grounded into coordinates by a grounding model
- Best combo (Gemini-3-Pro + UI-Inst-7B) averages 52%; end-to-end GUI-only models fare far worse
- The two new task axes are where models fail most
The author is building their own benchmark on top of this work.
More from coding & agent
- Build scripts are the blind spot of AI coding: wrong builds waste CI minutes and nearly shipped a broken release — jimmykoppel · 2026-09-07
- Hermes Desktop ships Mixture of Agents: multiple advisor models, one aggregator — alexcovo_eth · 2026-09-07
- Founder runs most of his business on an agentic software factory — hugobowne · 2026-09-07
- MIT's Jimmy Koppel: for any complex code, reading beats running it to understand behavior — jimmykoppel · 2026-09-07
- From code-built boat to detailed Blender ship: Astra's asset iteration — Dimillian · 2026-09-07
- Zero-code Redditor builds privacy-first local AI assistant using ChatGPT, Claude and Gemini CLI — Particular-Shape1972 · 2026-09-07