Noam Brown questions CAIS/Scale leaderboard's uniform 'reasoning high' setting

polynoamial · x · 2026-09-24

Noam Brown criticized the CAIS/Scale AI evaluation site's choice to run all models at "reasoning high", noting that the high setting varies wildly between models and makes comparisons questionable.

He suggested reporting the dollar cost of each evaluation, or better, plotting accuracy vs cost instead of reasoning effort levels.

Related event: Researchers question CAIS leaderboard's uniform 'reasoning high' benchmarking(2 posts)→

Original post →

More from Models

Models channel →