GSO benchmark near saturation; creator estimates ~5% of tasks in any benchmark are flawed
scaling01 · x · 2026-09-29
- slimshetty, creator of the GSO benchmark, says the leaderboard is close to saturation after a long gap in updates.
- He adds a candid methodological estimate: at least 5% of tasks in any benchmark — including his own — likely have issues.
- A first-hand note on the limits of current eval leaderboards.
More from Research
- Triangle Splatting SLAM: Imperial College's ECCV 2026 dense RGB-D SLAM with on-the-fly mesh extraction — rsasaki0109 · 2026-09-30
- Manifold opens early access: robotics eval platform runs thousands of GPU-parallel rollouts in 30 mins — paigeinsf · 2026-09-30
- 1,000 AI agents discover new CRISPR-like system in virus DNA within 24 hours — CurieuxExplorer · 2026-09-30
- Explaining just 5% of token positions retains nearly all audit success across 4.7M explanations — aisilab · 2026-09-30
- NTU's Persistence Forcing hits FID 1.63 on ImageNet 256 by heterogeneous refinement in pixel-space DiTs — NanyangTechnologicalUniversity · 2026-09-30
- IBM's Q&D trains proactive agents to ask better questions, beating a 15x larger model — ibm · 2026-09-30