New VLM paper finds image and text align from layer 1 in newer models
stanislavfort · x · 2026-07-22
A VLM paper challenges the common belief that image and text representations only align in the late layers.
- The authors used DeepDream-style optimization on Gemma 3.
- They found alignment appears from layer 1 in newer models, rather than only in the top layers.
- The post points to a paper co-authored by evzenwy, javirandor, floriantramer, and stanislavfort.
More from Research
- New paper links intelligence to a learnable-novelty view of Epiplexity — theomitsa · 2026-07-22
- Turing Motors says its CTO won gold in Kaggle’s 2026 ARC-AGI-linked contest — MeganRisdal · 2026-07-22
- Paper warns dubious Kaggle medical datasets are reaching both papers and clinics — EhudReiter · 2026-07-22
- Protein language models can learn homo-oligomer contacts from single sequences — anshulkundaje · 2026-07-22
- Commercial frontier models blocked attack forensics because they misread the responder — morqon · 2026-07-22
- LeCun’s JEPA pitch gets a concrete world-model paper behind it — nikola_mr64990 · 2026-07-22