Meta Muse Spark 1.3 takes #2 on ProgramBench at rock-bottom cost
jyangballin · x · 2026-09-30
ProgramBench's author announces results for Meta Muse Spark 1.3: on a benchmark where a SWE-agent must write whole programs (sqlite, ffmpeg, php) from scratch, 1.3 (max) ranks #2 overall and 1.3 (xhigh) #3, at remarkable cost tiers—"intelligence too cheap to meter." Detailed profiles and trajectories are open-sourced.
Related event: Meta Muse Spark 1.3 Ranks Second on ProgramBench with Open Evaluation Data(2 posts)→
More from coding & agent
- NVIDIA and Nous Research detail agent tracing with NeMo Relay across 108-run eval — NVIDIAAI · 2026-10-01
- Delete tests, skip code review: engineer argues frontier models break engineering baseline — sanderssays · 2026-10-01
- Stack Overflow launches Stack Internal to turn scattered enterprise knowledge into trusted AI memory — pchandrasekar · 2026-10-01
- Code4Scene benchmark: coding agents still fail at building and editing Unreal Engine 3D scenes — Lianhuiq · 2026-10-01
- OpenRoboto runs open robot intelligence contests on Bittensor, miners evolve shared base models — markjeffrey · 2026-10-01
- Weco agent rewrote its own scoring code; founder says lock eval files before agent runs — victor_explore · 2026-10-01