TabPFN-3.5 tops Kaggle's Otto competition out of the box, but experts call the benchmark flawed
RichmanRonald · x · 2026-09-16
TabPFN-3.5 launched today claiming rank 1 on the historic Kaggle Otto competition with no tuning — innixma argues this is more impressive than anything in the report, showing hidden signal in the data.
JFPuget pushed back: results on past Kaggle competitions mean little, for at least 3 reasons:
- You can overfit directly to the private leaderboard, something impossible for participants — destroying the whole point of measuring generalization;
- Winning solutions' code and write-ups are public;
- Today's models weren't available back then.
A representative debate on whether historical competition results make a valid benchmark for new models.
More from Models
- Jev is a smart classifier, not a frontier LLM — a contrarian analysis — DavideCrapis · 2026-09-16
- Jev debuts at $0.042/M input tokens with free output, claiming 20-200x speedup — hey_abusiddik · 2026-09-16
- Custom Midtrain Plus Own RL Matches Astra Max at Half the Inference Price — hsu_byron · 2026-09-16
- Astra Is Flawless for Hours, Then Randomly Stops and Makes Up Excuses — altryne · 2026-09-16
- Meta's new voice transcription model nearly 10x faster and more accurate than Whisper — alexandr_wang · 2026-09-16
- Qwen3.8 Max scores 45 on AA Intelligence Index, retakes China's top spot from GLM-5.3 — UmpireBorn3719 · 2026-09-16