MobileWorld benchmark: best phone GUI agents score only ~52% on 201 tasks

East-Muffin-6472 · reddit · 2026-08-19

The MobileWorld paper proposes a new benchmark for autonomous mobile agents: 201 tasks across 20 apps (comms, messaging, productivity, etc., <5% system apps), with two novel evaluation axes:

The setup uses a planner-executor architecture: a VLM planner sees only screenshots (no accessibility tree) and outputs natural-language actions, which a grounding model converts to precise (x,y) coordinates.

Results: the best combo, Gemini-3-Pro + UI-Inst-7B, averages 52%, end-to-end GUI models do far worse, and failures concentrate on the two new axes — proactive clarification and tool use remain clear weaknesses for phone agents.

Original post →

More from coding & agent

coding & agent channel →