CursorBench 4.0 launches; Muse Spark 1.3 matches Sol at under 40% the cost
jyangballin · x · 2026-09-11
leerob announced CursorBench 4.0, adding new tasks for instruction-following and sustained work on challenging projects, with a harder set that lowers all model scores. Quoting the launch, claireszhou highlighted that Muse Spark 1.3 matches Sol's performance at under 40% of the cost — at standard-tier pricing, not contributor pricing, which would widen the gap further on the y-axis.
More from coding & agent
- Bezalel launches one-URL agent platform bundling memory, email, money, texting, desktop, sandboxes and connectors — Rasmic · 2026-09-11
- Ten lessons from three years building agents for real production work — garrytan · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Shopify CEO Tobi Lütke hails single-dev open-source agent harness Pi — aakashgupta · 2026-09-11
- Dev explains why MCP won him over: organic UX beats telling agents to run CLI commands — zeeg · 2026-09-11
- Meta Details Muse Agent Safety: Sentinel Gate, Sandboxing and Human-in-the-Loop Protocol Approvals — altryne · 2026-09-11