OfirPress: benchmark gap comes from a 166-task subset and pass-rate vs full-completion metrics
OfirPress · x · 2026-09-23
OfirPress explains why another team reports much higher scores: they use a 166-task subset of his 200-task suite (likely dropping the 10-20 hardest), and report average total test pass rate rather than average task completion (tasks solved 100%). He notes models often only solve 60-70% of a task's tests, which inflates pass-rate metrics.
Related event: Ofir Press Explains Score Gap: Task Subsets and Scoring Methods(2 posts)→
More from Models
- Rumor: Anthropic's quiet next model; OpenAI reportedly near AGI, solved 100+ problems — brandon_galang · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Claude Opus 5.5 costs $5.98 per task as price cuts offset ~80% token usage spike — ArtificialAnlys · 2026-09-23
- Claude Opus 5.5 benchmarked: intelligence 58, per-task cost spans 11x across five tiers — ArtificialAnlys · 2026-09-23