AI researcher: benchmarks without released training data are 100% meaningless
mjdramstead · x · 2026-09-09
Researcher Michael Dramstead posted a blunt reminder: if a benchmark's training data isn't provided, the results are "absolutely, completely 100% meaningless," urging the community to stop pretending otherwise. The remark targets opaque train/test splits that make leaderboard scores incomparable and vulnerable to contamination.
More from Models
- GPT-6 Astra tops RSI-Exam at 0.5126, 18.4% above GPT-5.6 Sol — HuaxiuYaoML · 2026-09-09
- APEX-Agents 1.1 benchmark update: Claude Fable 5.1 tops leaderboard at 68.6% — amaarora · 2026-09-09
- Frontier labs must shrinkflate the $200 subscription to upsell you to API pricing — StewartalsopIII · 2026-09-09
- User math: peak pricing 2x but base rate better, 252M cache tokens cost just $0.75/day — teortaxesTex · 2026-09-09
- Astra noticeably worse than Sol in long threads, dev finds handoff workaround — jdjohnson · 2026-09-09
- VC communism is over: frontier models on rationing force hard model-choice thinking — StewartalsopIII · 2026-09-09