LeAVJEPA: one shared encoder for label-free audio-video self-supervised learning
arnosolin · x · 2026-10-08
The author introduces LeAVJEPA, extending LeJEPA (Balestriero & LeCun) to audio and video with a single shared transformer trained without labels.
- View construction: audio-only and video-only views learn to match a representation formed from both modalities
- Preventing shortcuts: if every view contained both modalities, matching could be solved via one modality alone; modality dropout makes the shortcut harder, with controlled ablations showing clear feature improvements
- Minimal architecture: embedding alignment unifies the modalities and LeJEPA's SIGReg regularizer prevents collapse — no EMA teacher, predictor, reconstruction decoder, or contrastive negatives needed
- Emergent behavior: audio-to-video attention follows sound-producing objects despite no localization supervision
- Results: with attentive probes on frozen ViT-L features, 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50
More from Research
- Hugging Face launches Robotic Episodes Viewer for 24k+ LeRobot datasets — mishig25 · 2026-10-09
- Blind humanoid walks, plays soccer and lifts suitcases with joint encoders only — accepted at Humanoids 2026 — Jan_R_Peters · 2026-10-09
- Delete object info from observations and PPO learns to search anyway — TU Darmstadt on its Humanoids 2026 paper — Jan_R_Peters · 2026-10-09
- U-Space finds an interpretable subspace for LLM uncertainty, no training needed — Tobias Braun · 2026-10-09
- CARE certifies VLA inference speedups up to 10.8x with statistical guarantees — UMCP · 2026-10-09
- SOL: a sample-based distributional metric proposed for evaluating text diffusion LMs — NandoDF · 2026-10-09