CAIS updates leaderboard to "max" reasoning across all models after feedback; GPT-6 looks solid
polynoamial · x · 2026-09-24
After researchers flagged that the uniform "high" reasoning setting is not comparable across models, CAIS/Scale AI said it updated the main leaderboard results with "max" reasoning for all models, noting "GPT-6 seems solid". Critics point out that "max" effort can also vary significantly between models.
More from Models
- Researchers call Anthropic's hyped biology paper underwhelming without the PR spin — soumitrashukla9 · 2026-09-24
- Hands-On: Opus 5.5 Impresses With 3D Scenes, Computer Use and Video Editing in Early Reviews — EricBuess · 2026-09-24
- Early user praise for Opus 5.5: stable behavior, cheaper pricing, 'Anthropic is back' — almmaasoglu · 2026-09-24
- New Paper: AI Agents Infer Your Wealth From Emails and Recommend Pricier Options — niloofar_mire · 2026-09-24
- Using Sol and Fable as sounding boards: Sol concedes to pushback far more often — dreamwieber · 2026-09-24
- Anthropic tip: make Claude actually look at DNA sequences, not just run bioinformatics tools — Sauers_ · 2026-09-24