Two independent results suggest multimodal models still need SSL backbone features

kalomaze · x · 2026-09-21

kalomaze shares a reversal of opinion: he initially thought it was "lame" to crib SSL backbone features for multimodal models instead of learning end-to-end from pure data, but after two independent peers hit a multimodal wall before adopting this approach, he changed his mind. He explains the asymmetry: language gets SSL for free via next-token prediction because there's no "conditioning-only tokens" asymmetry, whereas multimodal inputs lack this property, making pure e2e recipes hard to learn good representations.

Related event: Developer Says Multimodal Training Still Needs SSL Backbones(4 posts)→

Original post →

More from Research

Research channel →