Critic says CritPt and SciCode evals are broken, results untrustworthy until fixed

inductionheads · x · 2026-09-04

In a debate with TheZvi, teortaxesTex argues that both CritPt and SciCode benchmarks are broken, that there is genuine weakness in knowledge (and perhaps GDP-related tasks) versus Anthropic, and concludes results are untrustworthy until the evals are fixed.

Original post →

More from Models

Models channel →