COMODO distills video intelligence into IMU sensors for camera-free egocentric AI glasses

flosalim · x · 2026-10-12

Presented at UbiComp in Shanghai, COMODO is a self-supervised cross-modal distillation framework that transfers the semantic structure of pretrained video representations to IMU representations via similarity-distribution alignment. This enables egocentric human activity recognition with IMU-only inference, avoiding the battery drain, heat, and privacy issues of always-on cameras while preserving temporal structure. The thread also covers a companion study personalizing multimodal physiological sensing (EEG/fNIRS/ECG/EDA) for field depression screening across 194 users, quantifying the value of each additional label under different label budgets and meta-learning initializations.

Original post →

More from Embodied

Embodied channel →