SciCode benchmark outdated, >90% pass rates no longer reflect frontier

geoffwolfe · x · 2026-10-08

anshuman1 argues the SciCode-verified benchmark needs updating: >90% pass rates sound fine, but the problems are old and in some form already in training data. ReasonCoreAI is preparing newer, harder science problems and benchmarking against the latest model releases, aiming to build error-free evaluations that push the true frontier. Discussion arose in the context of congrats on a new Mistral open model's strong scores, including a leading cyber benchmark.

Original post →

More from Models

Models channel →