Tencent's AutoGUIWorld Uses Image Generators as World Models, Synthesizing 79K GUI Trajectories Without Running Software

Tencent-Hunyuan · hf · 2026-10-02

Tencent Hunyuan's AutoGUIWorld synthesizes GUI agent interaction trajectories without deploying software environments, combining image-generator visual priors with a planner's task knowledge. A planner specifies atomic actions and expected visual consequences while the image generator iteratively edits screenshots. Action grounding and transition-level filtering yield 79,266 spatially annotated step-level samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B lifts OSWorld mean score from 33.0% to 40.8% and ScienceBoard success rate from 14.0% to 32.2%.

Original post →

More from coding & agent

coding & agent channel →