ProgramBench opens the leaderboard to custom model-plus-harness submissions
jyangballin · x · 2026-07-25
Meta FAIR, Meta TBD, Stanford, and Harvard’s ProgramBench now lets users submit their own custom model + harness combinations to the leaderboard.
- The benchmark is expanding beyond fixed model-only submissions.
- A reply highlights that GLM 5.2 needed far more turns on average for a task than Opus 4.8.
- The key takeaway is that performance depends heavily on the harness, not just the model.
More from Models
- OpenAI and Hugging Face incident reignites debate over how scary misalignment really is — jammastergirish · 2026-07-25
- Anthropic’s rumored Opus 5 could ship with a 1 million token context window — imjustnewatai · 2026-07-25
- Kimi K3’s architecture is public, and the draft diagram shows KDA plus AttenRes — AccBalanced · 2026-07-25
- A user says $200/month Codex is so productive they now pay OpenAI $1,200 a month — robleclerc · 2026-07-25
- A Qwen3.6-27B merge blends reasoning and coding into one 27B model — pbaylies · 2026-07-25
- Google still indexes an Amazon Bedrock doc that mentions Claude Opus 5 — Angaisb_ · 2026-07-25