ICML Position Paper: Unlabeled Data Doesn't Mean No Human Supervision
serrjoa · x · 2026-09-05
An ICML 2026 position paper argues that 'unlabeled' does not mean 'supervision-free': data curation, preprocessing, and training objectives all embed human priors that models rely on.
The paper notes a telling trend — despite continued field growth, papers titled with 'unsupervised' at flagship vision conferences have sharply declined since 2021, signaling the umbrella term no longer captures real distinctions and makes cross-assumption comparisons unfair.
Recommendations:
- Authors should disclose priors in data selection and learning objectives
- Specify which pipeline components depend on which assumptions
- Standardized disclosure would improve communication, ensure fairer comparisons, and preserve methodological diversity
More from Research
- Homework for researchers: extending RoPE to tensor product representations — thomasahle · 2026-09-05
- Debate: Models Fuzzily Recall Concepts, Not Text — SAE Features vs Edit-Distance Memorization — voooooogel · 2026-09-05
- VLA-Corrector from ZJU & Alibaba DAMO lifts robot success rates while cutting policy calls — 机器之心 · 2026-09-05
- Spanda: Open-Source Hallucination Detector Runs in 1.5ms on CPU, 90,000x Faster than Semantic Entropy — Otherwise_Nobody_721 · 2026-09-05
- Bug Hunt Bench: 105 real bugs stress-test GPT-6, Claude, Grok, Gemini and more coding agents — PawelHuryn · 2026-09-05
- Many mathematicians value prestige over truth, discussion on AI proofs notes — avt_im · 2026-09-05