Noam Brown questions CAIS/Scale leaderboard's uniform 'reasoning high' setting
polynoamial · x · 2026-09-24
Noam Brown criticized the CAIS/Scale AI evaluation site's choice to run all models at "reasoning high", noting that the high setting varies wildly between models and makes comparisons questionable.
He suggested reporting the dollar cost of each evaluation, or better, plotting accuracy vs cost instead of reasoning effort levels.
More from Models
- Meta CAIO Alexander Wang: 'pretty soon we're dropping our most capable model ever' — ChrisGPT · 2026-09-24
- Meta's Alexandr Wang teases: 'pretty soon we are dropping the most capable model we have ever trained' — scaling01 · 2026-09-24
- Heavy user says new Claude Opus model is 'incredible' after days of pushing it hard — EricBuess · 2026-09-24
- Computer Use coming to Meta's Muse Mac app, teased at Meta Connect — kimmonismus · 2026-09-24
- Claude Opus 5.5 reportedly arrives with 40% cost cut and faster output — rohanpaul_ai · 2026-09-24
- Over-refusal pushes users to uncensored open models: 'working is less of a headache' — Khaledthe · 2026-09-24