Researcher challenges CAIS leaderboard's uniform "high" reasoning setting, urges cost reporting
polynoamial · x · 2026-09-24
Researcher polynoamial publicly questioned how CAIS/Scale AI evaluates all models at the uniform "reasoning high" setting. He argues that "high" reasoning effort varies widely across models, so a single setting doesn't guarantee comparability, and suggests reporting the dollar cost of each evaluation or plotting accuracy vs. cost instead of a single score.
More from Models
- Developer says Claude usage shifted from Fable limits to mostly Opus limits this week — rudrank · 2026-09-24
- Users call out OpenAI: Sol looks like rebranded Terra with $2 cheaper output but worse usage limits — RexDouglass · 2026-09-24
- That Opus 5.5 Demoscene Demo Used Under 3% of a Weekly Usage Quota — gandamu_ml · 2026-09-24
- Researchers find Mythos activations behave like genomic language models, hinting LLMs learn DNA motifs — jatin_n0 · 2026-09-24
- GLM-5.3 crushes stealth model Space Bunny Alpha on Newton's cradle physics test — rohanpaul_ai · 2026-09-24
- theo: Anthropic has no small models worth using, OpenAI no large ones, Google none at all — james_mtc · 2026-09-24