ProgramBench adds Pareto curves as models improve sharply just two months after launch
jyangballin · x · 2026-07-27
ProgramBench now includes Pareto curves for per-instance analysis, making it easier to inspect the trade-off between cost and performance.
The author notes that:
- the new curves help visualize progress at the instance level,
- the benchmark uses % tests passed for illustration,
- but % resolved is still the recommended official metric.
The attached chart shows strong movement on cost/performance fronts only two months after release, suggesting models are improving quickly on this benchmark and should be evaluated sooner rather than later.
Related event: ProgramBench Adds Pareto Curves to Track Rapid Model Progress(3 posts)→
More from Models
- Nebius readies for tomorrow’s Kimi launch, with a bigger team and hiring stakes — demian_ai · 2026-07-27
- Claude Opus 5 ranks third on VoxelBench, just 30 Elo behind the leader — legit_api · 2026-07-27
- Open source and open weights are not the same thing, a post argues — RisingSayak · 2026-07-27
- Open-source model debates are turning tribal, and confidence is outrunning evidence — matt_slotnick · 2026-07-27
- Moonshot teases Kimi-K3 with a July 27, 2026 release countdown — teortaxesTex · 2026-07-27
- Codex users report higher credit burn as GPT-5.6 makes sequential tool calls — kevinkern · 2026-07-27