Researcher challenges CAIS leaderboard's uniform "high" reasoning setting, urges cost reporting

polynoamial · x · 2026-09-24

Researcher polynoamial publicly questioned how CAIS/Scale AI evaluates all models at the uniform "reasoning high" setting. He argues that "high" reasoning effort varies widely across models, so a single setting doesn't guarantee comparability, and suggests reporting the dollar cost of each evaluation or plotting accuracy vs. cost instead of a single score.

Related event: Researchers question CAIS leaderboard's uniform 'reasoning high' benchmarking(2 posts)→

Original post →

More from Models

Models channel →