OfirPress: benchmark gap comes from a 166-task subset and pass-rate vs full-completion metrics

OfirPress · x · 2026-09-23

OfirPress explains why another team reports much higher scores: they use a 166-task subset of his 200-task suite (likely dropping the 10-20 hardest), and report average total test pass rate rather than average task completion (tasks solved 100%). He notes models often only solve 60-70% of a task's tests, which inflates pass-rate metrics.

Related event: Ofir Press Explains Score Gap: Task Subsets and Scoring Methods(2 posts)→

Original post →

More from Models

Models channel →