Tencent's AutoGUIWorld Uses Image Generators as World Models, Synthesizing 79K GUI Trajectories Without Running Software
Tencent-Hunyuan · hf · 2026-10-02
Tencent Hunyuan's AutoGUIWorld synthesizes GUI agent interaction trajectories without deploying software environments, combining image-generator visual priors with a planner's task knowledge. A planner specifies atomic actions and expected visual consequences while the image generator iteratively edits screenshots. Action grounding and transition-level filtering yield 79,266 spatially annotated step-level samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B lifts OSWorld mean score from 33.0% to 40.8% and ScienceBoard success rate from 14.0% to 32.2%.
More from coding & agent
- Dev recreates a DOOM-like game with GPT, nailing the original's gory feel — DeryaTR_ · 2026-10-02
- Process-mining agents found 20 steps and 7 loops in a workflow documented as 7 steps — vasuman · 2026-10-02
- Geoffrey Huntley: Forget reading code—your verification properties are all that matters — kieranklaassen · 2026-10-02
- Claude Skills explained: why prewritten PDF scripts beat pasting prompts every time — lxfater · 2026-10-02
- Eye.Art Polyphemus: a chat-first MCP for image generation and reference-based edits — axiomofaxiom · 2026-10-02
- Dev uses GitHub Copilot as a project lead: files issues, lets the agent do everything else — 0xkarasy · 2026-10-02