ByteDance's WorkflowGym Exposes GUI Agents: Top Models Pass Only 30%

机器之心 · wechat · 2026-07-23

ByteDance's Seed evaluation team has released WorkflowGym, a new benchmark that tests GUI agents on professional software like FreeCAD and Blender, revealing the true boundaries of current model capabilities.

Evaluation Design & Results:

Core Model Flaws:

The paper identifies four major dead spots in current agent frameworks: loss of long-context consistency, cascading errors without self-correction, goal drift, and a lack of understanding of professional UI paradigms. This suggests that scaling parameters alone won't fix these architectural bottlenecks, making human-AI collaboration the most practical product approach for now.

Original post →

More from Models

Models channel →