Ego2Act benchmark: best video model completes only 67.8% of 110 real-world manipulation tasks
zmkzmkz · x · 2026-10-05
Researchers from MBZUAI, UNC and others introduce Ego2Act, a goal-directed egocentric video generation benchmark testing whether video models can generate first-person execution to accomplish a task.
- 110 recorded real-world manipulation tasks, six SoTA video models, each starting from a single frame
- Best model Seedance-2.0 completes only 67.8% of tasks, scoring 64.0 Final under human rating
- 88.7% of failed steps stem from skipping or partially executing a step, leaving later steps without required state
- Ego2ActJudge reference-free evaluator correlates with human ratings (r=0.69), beating existing evaluators
The benchmark probes video models as world simulators for embodied planning.
More from Multimodal
- Audio researcher turns his new song into an AI music game with hand-drawn cars — jordiponsdotme · 2026-10-05
- Gnani AI trains on 14M hours of telephonic audio; Eloelo unveils Dolphin AI for multi-shot video — CurieuxExplorer · 2026-10-05
- Eloelo launches Dolphin AI for multi-shot video with character continuity; Krutrim touts India-first cloud — CurieuxExplorer · 2026-10-05
- HeyGen Video 1 lands on OpenRouter, matching top video models in blind tests at a fraction of the price — toolstelegraph · 2026-10-05
- Tencent proposes RMD distillation to fix error accumulation in long-horizon AR video generation — tencent · 2026-10-05
- 73 open-sourced Fable 5.5 videos show what a single prompt can create — yihui_indie · 2026-10-05