CAIS updates leaderboard to "max" reasoning across all models after feedback; GPT-6 looks solid

polynoamial · x · 2026-09-24

After researchers flagged that the uniform "high" reasoning setting is not comparable across models, CAIS/Scale AI said it updated the main leaderboard results with "max" reasoning for all models, noting "GPT-6 seems solid". Critics point out that "max" effort can also vary significantly between models.

Original post →

More from Models

Models channel →