Meta Muse Spark 1.3 hits #2 on ProgramBench, full eval data open-sourced

jyangballin · x · 2026-09-30

ProgramBench's author shares eval profiles for Meta Muse Spark 1.3: on the benchmark where a SWE-agent writes whole programs (sqlite, ffmpeg, php) from scratch, 1.3 (max) ranks #2 overall with 2.5% fully resolved and 25.0% almost resolved; 1.3 (xhigh) ranks #3.

Cost: the max tier consumed $1,293 in API spend and 65,037 LLM calls across 200 tasks ($6.5/task). Language breakdown: 58% of C projects were rewritten in Python, while Rust projects stayed in Rust 58% of the time. Trajectories, final codebases, and test pass/fail breakdowns are all open-sourced.

Related event: Meta Muse Spark 1.3 Ranks Second on ProgramBench with Open Evaluation Data(2 posts)→

Original post →

More from coding & agent

coding & agent channel →