CursorBench 4.0 launches; Muse Spark 1.3 matches Sol at under 40% the cost

jyangballin · x · 2026-09-11

leerob announced CursorBench 4.0, adding new tasks for instruction-following and sustained work on challenging projects, with a harder set that lowers all model scores. Quoting the launch, claireszhou highlighted that Muse Spark 1.3 matches Sol's performance at under 40% of the cost — at standard-tier pricing, not contributor pricing, which would widen the gap further on the y-axis.

Related event: CursorBench 4.0 Rolls Out Harder Long-Horizon Tasks, Model Scores Drop Across the Board(2 posts)→

Original post →

More from coding & agent

coding & agent channel →