Yoav Goldberg: models aren't overfitting tests, they're trained for these tasks at scale
yoavgo · x · 2026-09-14
Researcher Yoav Goldberg clarifies his position in a debate: he didn't claim models "overfit to tests," but that they were "trained for the task," likely at very large scale. He still thinks they may generalize well to such tasks, amid ongoing arguments over whether benchmark performance reflects real capability.
More from Models
- Agent trace dataset hits 50k+ monthly downloads — author speculates on SFT and reward-hacking monitor uses — maksym_andr · 2026-09-14
- Author Finds Gemini 'Reliably Wrong' at Verifying Quote Sources, 0/2 in Tests — danbri · 2026-09-14
- Astrable: Open-Source Codex Plugin Pairs GPT-6 Astra with Claude Fable 5.1 — daniel_mac8 · 2026-09-14
- GPT-Live-1 tested on real phone calls: natural speech but serious instruction-following flaws — kolchinski · 2026-09-14
- Dan Shipper: using Astra medium for simple tasks, "pacing the frontier" — danshipper · 2026-09-14
- Grok-4.6 spotted generalizing well across ARC-AGI generations, per third-party observation — zainhas · 2026-09-14