Meta, Stanford and Harvard open up ProgramBench leaderboard with community submissions
jyangballin · x · 2026-09-30
ProgramBench—a joint effort across Meta Superintelligence Labs, Stanford, and Harvard—launched a community leaderboard and open-sourced everything: trajectories, final codebases, and test pass/fail breakdowns, with eval profiles and generated codebases for Muse Spark 1.3 already public.
- How to submit: run evaluations with uvx programbench (1.2.0+), package your run dir (submission.yaml, per-task trajectories and eval results) into a public GitHub repo, optionally offloading heavy artifacts to a HuggingFace dataset;
- A registry PR to ProgramBench/submissions puts your row on the leaderboard once merged;
- The official Gemini 3.1 Pro + mini-SWE-agent baseline serves as the example submission;
- More models are coming soon—6 astra and 5.5 opus—and submissions for your own model or harness are open now.
More from coding & agent
- Delete tests, skip code review: engineer argues frontier models break engineering baseline — sanderssays · 2026-10-01
- Stack Overflow launches Stack Internal to turn scattered enterprise knowledge into trusted AI memory — pchandrasekar · 2026-10-01
- Code4Scene benchmark: coding agents still fail at building and editing Unreal Engine 3D scenes — Lianhuiq · 2026-10-01
- OpenRoboto runs open robot intelligence contests on Bittensor, miners evolve shared base models — markjeffrey · 2026-10-01
- Weco agent rewrote its own scoring code; founder says lock eval files before agent runs — victor_explore · 2026-10-01
- Real-world agent finance ops: escalates EUR 9,000 bill over approval limit, avoids duplicate payments — kimmonismus · 2026-10-01