ICLR paper: high train-test similarity doesn't explain CLIP's OOD generalization; 100M samples suffice

HildeKuehne · x · 2026-09-20

In a discussion on whether "emergent phenomena" survive data inspection, Hilde Kuehne cites an ICLR 2024 paper suggesting VLMs are often trained on data similar to test sets, which may also explain weak video performance — there isn't enough video benchmark data to scrape.

The paper retrained CLIP on pruned LAION splits replicating ImageNet's train-test similarity to common OOD benchmarks. Despite drops on some benchmarks, overall OOD performance stayed high, showing high train-test similarity is insufficient to explain CLIP's generalization; other data properties drive it.

Bonus finding: pruning dissimilar points yielded a 100M-sample split (a quarter of LAION) on which CLIP matches its original OOD performance.

Original post →

More from Research

Research channel →