Two independent results suggest multimodal models still need SSL backbone features
kalomaze · x · 2026-09-21
kalomaze shares a reversal of opinion: he initially thought it was "lame" to crib SSL backbone features for multimodal models instead of learning end-to-end from pure data, but after two independent peers hit a multimodal wall before adopting this approach, he changed his mind. He explains the asymmetry: language gets SSL for free via next-token prediction because there's no "conditioning-only tokens" asymmetry, whereas multimodal inputs lack this property, making pure e2e recipes hard to learn good representations.
Related event: Developer Says Multimodal Training Still Needs SSL Backbones(4 posts)→
More from Research
- 'Math is not yet ready': Collatz conjecture may yield to AI in coming years — burny_tech · 2026-09-21
- Reverse-engineering Jev: its architecture, philosophy, and where it falls down — iamrobotbear · 2026-09-21
- Honglak Lee at Nobel Prize Dialogue: AI accelerates science, but skepticism matters more — honglaklee · 2026-09-21
- HF's Merve Noyan shares beginner guides for zero-shot classifiers, says many LLM tasks never needed them — ceciletamura · 2026-09-21
- Researchers formally verify the Kubernetes control plane with a compositional CORE spec — tianyin_xu · 2026-09-21
- Astra shows any 3D/4D prior can be distilled into VLMs, a new embodied AI paradigm — mariyaivasileva · 2026-09-21