Ofir Press explains why third-party agent benchmark scores look higher: subset + different metric

OfirPress · x · 2026-09-23

Asked why third parties report much higher scores than his official leaderboard, Ofir Press explained two reasons: (1) they use a subset of 166 tasks, likely discarding the 10-20 super-hard ones from the full 200-task set; (2) they report average total test pass rate, while his leaderboard reports average task completion (tasks solved 100%). A useful reminder that task selection and metric definitions can significantly inflate benchmark comparisons.

Related event: Ofir Press Explains Score Gap: Task Subsets and Scoring Methods(2 posts)→

Original post →

More from Models

Models channel →