COMODO distills video intelligence into IMU sensors for camera-free egocentric AI glasses
flosalim · x · 2026-10-12
Presented at UbiComp in Shanghai, COMODO is a self-supervised cross-modal distillation framework that transfers the semantic structure of pretrained video representations to IMU representations via similarity-distribution alignment. This enables egocentric human activity recognition with IMU-only inference, avoiding the battery drain, heat, and privacy issues of always-on cameras while preserving temporal structure. The thread also covers a companion study personalizing multimodal physiological sensing (EEG/fNIRS/ECG/EDA) for field depression screening across 194 users, quantifying the value of each additional label under different label budgets and meta-learning initializations.
More from Embodied
- Unitree open-sources 6B humanoid model UnifoLM-WLA-1.0, claims fine-tuning on a 24GB GPU — lmoroney · 2026-10-12
- Tesla Cybercab tracker logs 145 on the road, 319 registered with Texas DMV — JOBhakdi · 2026-10-12
- Liquid AI open-sources Open d1 decision models: d1-3B runs at 35ms per image on Jetson Thor — TheTuringPost · 2026-10-12
- Fleet Week traffic trapped Andrew Chen's Waymo for 15 min until humans-in-the-loop rerouted it — andrewchen · 2026-10-12
- DirtyMoCap recovers 3D motion from noisy unordered markers, with 100x faster CUDA solver — jamestagg · 2026-10-12
- AI agent deploys a micro factory end-to-end: €20k cost, ROI under 2 years — ihorbeaver · 2026-10-12