Ofir Press Explains Score Gap: Task Subsets and Scoring Methods
Ofir Press explained why third-party benchmark scores exceed his official leaderboard: they used only 166 of 200 tasks, likely dropping the hardest 10-20, and averaged pass rates in a way that masks true task completion.
2026-09-23 ~ 2026-09-23 · 2 related posts
- Ofir Press explains why third-party agent benchmark scores look higher: subset + different metric — OfirPress · 2026-09-23
- OfirPress: benchmark gap comes from a 166-task subset and pass-rate vs full-completion metrics — OfirPress · 2026-09-23