Ego2Act benchmark: best video model completes only 67.8% of 110 real-world manipulation tasks

zmkzmkz · x · 2026-10-05

Researchers from MBZUAI, UNC and others introduce Ego2Act, a goal-directed egocentric video generation benchmark testing whether video models can generate first-person execution to accomplish a task.

The benchmark probes video models as world simulators for embodied planning.

Original post →

More from Multimodal

Multimodal channel →