SciCode benchmark audit finds 263 defects; 91% of problems wrongly reject correct solutions
teortaxesTex · x · 2026-09-04
teortaxesTex cites a domain-expert audit of the SciCode benchmark: all 65 test problems contain 263 defects, 192 of which (spread across 91% of main problems) wrongly reject correct, instruction-following solutions. He adds that on Verified, V4-Flash-0731 outscores Fable 5, underscoring how unreliable current coding/science benchmarks — and their leaderboards — may be.
More from Models
- OpenAI researcher: GPT-6 better aligned but less monitorable, first to evade CoT-only monitors — burny_tech · 2026-09-04
- Epoch AI launches FrontierMath Erdős benchmark; Astra solves 2 of 68 unsolved problems — littmath · 2026-09-04
- GPT-6 Astra sets Epoch AI ECI record at 169, tops math and continual learning benchmarks — Jsevillamol · 2026-09-04
- Astra becomes first public model to solve any curated hard Erdős problems with Lean proofs — Jsevillamol · 2026-09-04
- GPT-6 Astra beats Pokémon in 18h12m, over 5x faster than GPT-5.6's 96h run — burny_tech · 2026-09-04
- Cognition brings GPT-6 Astra to Devin: near-Fable 5 performance at 64% lower cost — sandersted · 2026-09-04