MBZUAI's Ego2Act Benchmark Shows Video Models Skip Steps and Break Physical Plausibility
MBZUAI · hf · 2026-10-03
MBZUAI introduces Ego2Act, a benchmark for goal-directed egocentric video generation testing video models as world simulators:
- Scale: 2,640 videos across 110 real-world everyday tasks with varying clutter and multi-step complexity; given a scene image and high-level goal, models must simulate hands manipulating objects to complete the task.
- Automated eval: Ego2ActJudge, a reference-free judging pipeline, aligns with human consensus better than baselines.
- Key findings: generated simulations often skip or partially execute steps, leaving later steps missing dependent states and goals unfulfilled; models consistently fail at fine-grained physics, complex manipulation, and persistent world modeling.
The authors position Ego2Act as a rigorous testbed for physically plausible, goal-directed simulation.
More from Embodied
- MolmoMotion: AI2's 4B VLM forecasts 3D point trajectories from language instructions — rsasaki0109 · 2026-10-03
- Video shows Waymo robotaxi navigating city streets and merging onto freeway — reed · 2026-10-03
- Watching My Home Robot Rooma Get Lost for the 1000th Time — OnlineInference · 2026-10-03
- OpenAI's Dot hardware is legit despite a broken demo, early review says — altryne · 2026-10-03
- ALOHA creator quits Stanford PhD for Sunday Robotics; LeRobot's Cadene builds UMA — FinanceYF5 · 2026-10-03
- 5 accounts to follow for a ~6-month head start on robotics progress — FinanceYF5 · 2026-10-03