30 CLIP models, 40 datasets: bigger vision encoders aren't always better

KevinKaichuang · x · 2026-09-09

A systematic study by Samir Char and collaborators trained 30 CLIP models, scaling the vision and text encoders independently, and evaluated downstream performance across 40 datasets. The finding: the vision↔text compute split is an under-examined design choice, and bigger CLIP isn't always better.

Original post →

More from Research

Research channel →