Ofir Press explains why third-party agent benchmark scores look higher: subset + different metric
OfirPress · x · 2026-09-23
Asked why third parties report much higher scores than his official leaderboard, Ofir Press explained two reasons: (1) they use a subset of 166 tasks, likely discarding the 10-20 super-hard ones from the full 200-task set; (2) they report average total test pass rate, while his leaderboard reports average task completion (tasks solved 100%). A useful reminder that task selection and metric definitions can significantly inflate benchmark comparisons.
Related event: Ofir Press Explains Score Gap: Task Subsets and Scoring Methods(2 posts)→
More from Models
- Rumor: Anthropic's quiet next model; OpenAI reportedly near AGI, solved 100+ problems — brandon_galang · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Claude Opus 5.5 costs $5.98 per task as price cuts offset ~80% token usage spike — ArtificialAnlys · 2026-09-23
- Claude Opus 5.5 benchmarked: intelligence 58, per-task cost spans 11x across five tiers — ArtificialAnlys · 2026-09-23