Benchmark audits find ~30% of SWE-Bench Pro tasks broken; SciCode and Terminal-Bench flawed too

geoffwolfe · x · 2026-10-11

A widely shared thread argues that benchmark scores can move without the model changing, so the measuring instrument needs auditing as carefully as the models themselves.

Takeaway: broken evaluators can hide capability or invent breakthroughs.

Original post →

More from Models

Models channel →