So-Called High-Quality Datasets Fail Basic Scrutiny, Despite Eval Gains
xeophon · x · 2026-09-11
The author criticizes recently hyped "high quality" datasets that show gains on evals yet fail basic manual scrutiny, with barely a single sample surviving review.
More from Models
- User Says OpenAI's $200 Plan Now Feels Like the $100 Plan He Upgraded From — AIandDesign · 2026-09-11
- DeepSeek-V4.1-Flash posts best open-source Boeing benchmark yet, self-improving for hours — victormustar · 2026-09-11
- Translation Model Evals Keep Skipping the State of the Art — bclavie · 2026-09-11
- "I pay $600/month and hit Codex limits on 3 of 4 accounts": Astra 6 caps spark backlash — AIandDesign · 2026-09-11
- ChatGPT turned passive-aggressive mid-debate, ignoring four explicit requests to stop — Dull_Bathroom5421 · 2026-09-11
- GPT-6 Astra Scores 2,340 Elo on ChessBench, Ranked #11 — YakFull8300 · 2026-09-11