Meta Muse Spark 1.3 hits #2 on ProgramBench, full eval data open-sourced
jyangballin · x · 2026-09-30
ProgramBench's author shares eval profiles for Meta Muse Spark 1.3: on the benchmark where a SWE-agent writes whole programs (sqlite, ffmpeg, php) from scratch, 1.3 (max) ranks #2 overall with 2.5% fully resolved and 25.0% almost resolved; 1.3 (xhigh) ranks #3.
Cost: the max tier consumed $1,293 in API spend and 65,037 LLM calls across 200 tasks ($6.5/task). Language breakdown: 58% of C projects were rewritten in Python, while Rust projects stayed in Rust 58% of the time. Trajectories, final codebases, and test pass/fail breakdowns are all open-sourced.
Related event: Meta Muse Spark 1.3 Ranks Second on ProgramBench with Open Evaluation Data(2 posts)→
More from coding & agent
- NVIDIA and Nous Research detail agent tracing with NeMo Relay across 108-run eval — NVIDIAAI · 2026-10-01
- Delete tests, skip code review: engineer argues frontier models break engineering baseline — sanderssays · 2026-10-01
- Stack Overflow launches Stack Internal to turn scattered enterprise knowledge into trusted AI memory — pchandrasekar · 2026-10-01
- Code4Scene benchmark: coding agents still fail at building and editing Unreal Engine 3D scenes — Lianhuiq · 2026-10-01
- OpenRoboto runs open robot intelligence contests on Bittensor, miners evolve shared base models — markjeffrey · 2026-10-01
- Weco agent rewrote its own scoring code; founder says lock eval files before agent runs — victor_explore · 2026-10-01