Expert Warns: Kaggle Public Leaderboard Scores Prone to Overfitting
calabi_and_yau · x · 2026-08-13
In response to recent claims that the AI model 'Locus' outperformed 89.5% of human competitors across multiple Kaggle competitions, expert Jean-François Puget raised strong skepticism.
He pointed out that the ongoing competitions' current scores are solely based on the public leaderboard. Climbing the public LB can often be achieved through aggressive overfitting, which does not reflect the model's actual performance on the final hidden test data. Furthermore, the claimed '4th place average' is misleading because it only looks at the few humans who enter all competitions, who are typically not the top contestants.
More from Models
- Qwen3.8 release imminent: users speculate on improvements, jokingly request 'monk mode' reasoning — Ok-Shower7286 · 2026-08-13
- Qwen3.8-27B Model Countdown Page Goes Live on Hugging Face — paf1138 · 2026-08-13
- Anthropic Overtakes OpenAI in Enterprise Adoption, Ramp Data Shows — rohanpaul_ai · 2026-08-13
- Qwen-Max Benchmark Scores Fluctuate: Drops to 53 on First Run — teortaxesTex · 2026-08-13
- xAI Accused of Omitting Safety and Prompt Injection Robustness Results — npinto · 2026-08-13
- Speculation Suggests Anthropic Might Be Hiding a Claude 3.5 Pro Model — teortaxesTex · 2026-08-13