ICLR paper: high train-test similarity doesn't explain CLIP's OOD generalization; 100M samples suffice
HildeKuehne · x · 2026-09-20
In a discussion on whether "emergent phenomena" survive data inspection, Hilde Kuehne cites an ICLR 2024 paper suggesting VLMs are often trained on data similar to test sets, which may also explain weak video performance — there isn't enough video benchmark data to scrape.
The paper retrained CLIP on pruned LAION splits replicating ImageNet's train-test similarity to common OOD benchmarks. Despite drops on some benchmarks, overall OOD performance stayed high, showing high train-test similarity is insufficient to explain CLIP's generalization; other data properties drive it.
Bonus finding: pruning dissimilar points yielded a 100M-sample split (a quarter of LAION) on which CLIP matches its original OOD performance.
More from Research
- WindTunnel benchmark: WebMCP makes browser agents 2.5-7.5x faster, 3-47x cheaper — FinanceYF5 · 2026-09-20
- Gary Marcus disputes LLM 'pain direction' paper: language clusters don't mean suffering — anilkseth · 2026-09-20
- MatSemNet uses LLMs to mine 700+ papers, modeling reaction pathways as sequences for catalyst discovery — bravo_abad · 2026-09-20
- Smart Zoi runs 50 AI NPCs in real time with a fine-tuned 1B SLM, LIFT hits 100+ citations — Kangwook_Lee · 2026-09-20
- LlamaIndex benchmarks 100+ models on document parsing with ParseBench — solyarisoftware · 2026-09-20
- ICLR gets record submissions; authors' journal spam emails hit records too — wandedob · 2026-09-20