LeAVJEPA: minimalist audio-visual SSL hits 36.0 mAP on AudioSet and 91.3% on ESC-50 frozen
arnosolin · x · 2026-10-08
Arno Solin's group at Aalto University and ELLIS Institute Finland released LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective.
- Minimalist architecture: a single early-fusion Vision Transformer handles audio, video, and joint audio-video inputs — no EMA target encoders, prediction heads, reconstruction decoders, or contrastive losses.
- Key mechanism: modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit; SIGReg prevents collapse. Ablations identify modality dropout as the crucial ingredient for alignment.
- Emergent localization: audio-to-video attention follows sound-producing objects with zero localization supervision.
- Results: with attentive probes on frozen ViT-L features, 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50; 61.1% on VGGSound after fine-tuning, plus zero-shot audio-visual retrieval.
Paper, project page (with interactive demos), and code are available.
More from Research
- PersistBench wins NeurIPS Spotlight, finds 4D foundation models lack visual memory — weichiuma · 2026-10-09
- StarkWare founder Eli Ben-Sasson: AI solved the Erdős Unit Distance problem, all bets are off — jamestagg · 2026-10-09
- Claude Science produces first complete ultraviolet map of the sky, ~10% deviation — The Decoder · 2026-10-09
- FreeMatching: generalizable dense correspondence matching beyond spatio-temporal priors — hkuhk · 2026-10-09
- MBZUAI's WorldGuide beats MiniMax-H3 on closed-loop procedural video world modeling — MBZUAI · 2026-10-09
- REMORY adds soft residual memory tokens to context compaction, hitting near full-context scores at 5.2% of input — Hanchen Xia · 2026-10-09