Tabular learning needs cross-validation across hundreds of tables, researcher says
GaelVaroquaux · x · 2026-07-26
Gael Varoquaux says tabular learning evaluation is still badly done: serious evaluation requires cross-validation across hundreds of tables, each with thousands or even hundreds of thousands of rows.
He argues that past failures to evaluate properly have contributed to stagnation in the field, and that model size alone is not the right question. The reply thread jokes about how large a model would satisfy reviewers, but the core point is that rigorous evaluation methodology matters more than chasing scale.
Related event: Researchers Discuss PFNs vs Post-trained LLMs for Tabular Data(4 posts)→
More from Research
- Orca opens a shared stack for dexterous hand research and teleoperation — tensorqt · 2026-07-26
- Weekly AI paper roundup covers world models, embodied AI, and self-improving agents — _akhaliq · 2026-07-26
- PFNs train from scratch for tabular data, while post-trained LLMs hit limits fast — roydanroy · 2026-07-26
- Paper finds RL gains can be predicted from pretraining loss, with optimal compute share around 20%–28% — rohanpaul_ai · 2026-07-26
- A 2007 survey on collective communication still helps when learning NCCL — abhi9u · 2026-07-26
- Alibaba and CUHK’s RynnWorld-4D predicts a robot scene’s future from one image — jiqizhixin · 2026-07-26