ECCV talk outlines three pillars for embodied AI: motion prediction, evidence, streaming

CSProfKGD · x · 2026-09-10

Juan Carlos Niebles' ECCV 2026 CONTEXTUS workshop talk argues embodied AI systems need three foundational capabilities: predictive motion understanding via UniEgoMotion (ICCV 2025, anticipating human action from egocentric video), evidence-backed reasoning via E-VQA (dense spatio-temporal grounding so systems show their work), and streaming efficiency via StateKV (real-time inference over hours of continuous video). Slides are publicly available.

Original post →

More from Embodied

Embodied channel →