ECCV 2026 tutorial tackles the evaluation bottleneck for visual foundation and world models

georgiagkioxari · x · 2026-09-08

Ziqi Ma will present StEvo and new findings on evaluating world models at ECCV 2026's tutorial "Evaluating Visual Foundation and World Models" (Sept 9, Malmö).

The tutorial's core thesis: as visual AI matures into interactive, 3D-consistent, action-conditioned simulation, evaluation becomes the primary bottleneck — outputs must not just look right but stay coherent over time, obey physics, and support agent interaction.

Covers three levels: visual understanding benchmarks (hallucination, robustness, contamination); image/video generation (compositionality, physical plausibility, reward hacking risks with learned reward models); and world models (3D consistency, geometric stability, state evolution). Related paper on state evolution in video world models shows at the Sept 11 poster session.

Original post →

More from Research

Research channel →