Researchers question CAIS leaderboard's uniform 'reasoning high' benchmarking

Noam Brown and others criticized CAIS/Scale AI's leaderboard for benchmarking all models at a uniform 'reasoning high' setting, arguing the setting means different things across models, making results incomparable and costs opaque.

2026-09-24 ~ 2026-09-24 · 2 related posts