SciCode benchmark outdated, >90% pass rates no longer reflect frontier
geoffwolfe · x · 2026-10-08
anshuman1 argues the SciCode-verified benchmark needs updating: >90% pass rates sound fine, but the problems are old and in some form already in training data. ReasonCoreAI is preparing newer, harder science problems and benchmarking against the latest model releases, aiming to build error-free evaluations that push the true frontier. Discussion arose in the context of congrats on a new Mistral open model's strong scores, including a leading cyber benchmark.
More from Models
- ThePrimeagen benchmarks OpenAI's new decision model: faster and more accurate on GUI tasks — prd_008 · 2026-10-08
- ~90% of frontier lab compute now goes to post-training and inference: Social Capital — soumitrashukla9 · 2026-10-08
- AI2's Bolmo tackles the 'token tax' hitting Global South scripts — Kyle_L_Wiggers · 2026-10-08
- Grok bot goes proactive on X, now suggests what you should do next — Daniel_Farinax · 2026-10-08
- OpenAI and Cloudflare launch decision APIs — tested at 355 decisions, TypeSafe's Jev still decides more — PawelHuryn · 2026-10-08
- Gary Marcus: OpenAI's vague math report 'would never pass peer review' — Tao responds too — Gary Marcus · 2026-10-08