30 CLIP models, 40 datasets: bigger vision encoders aren't always better
KevinKaichuang · x · 2026-09-09
A systematic study by Samir Char and collaborators trained 30 CLIP models, scaling the vision and text encoders independently, and evaluated downstream performance across 40 datasets. The finding: the vision↔text compute split is an under-examined design choice, and bigger CLIP isn't always better.
More from Research
- New EMNLP paper uses controlled rearing to explain how learners reject ungrammatical "X laughed Y" — EliasEskin · 2026-09-10
- Greg Kamradt digs into gamedevbench: reference solutions edit 4.7 files, 114 lines on average — GregKamradt · 2026-09-10
- Ex-Continue founder publishes 'The Superdark Factory' in MIT Press: 35T agent tokens/month by 2029 — tylerjdunn · 2026-09-10
- Experiment with 100 AI agents: when 9% cheated on math problems, 24% blew the whistle — weballergy · 2026-09-10
- Scientific Imaging and ML workshop set for San Juan, March 2027 — applications open until Oct 2026 — prof_kamilov · 2026-09-10
- Lesioning one neuron in a fruit fly brain model makes it misidentify emotions — emax · 2026-09-10