SWE-Together Update: Claude Fable 5 Tops Coding Benchmark, Muse Spark 1.3 Is 5x Cheaper
shuchaobi · x · 2026-09-11
SWE-Together, a benchmark of 109 real coding tasks with a simulated user in the loop, published an updated leaderboard with key findings:
- Claude Fable 5 leads at 61% judge score, with Fable 5.1 at 57%. Peak performance is nearly identical; the gap is consistency — Fable 5.1 scored zero on 14 trials vs 6 for Fable 5. But 5.1 is 25% faster, 2x cheaper, and needs slightly fewer user corrections (1.45 vs 1.53 per task).
- Meta's Muse Spark 1.3 is the value outlier: 2.5x cheaper per solved task than Fable 5.1 and 5x cheaper than Fable 5, with mid-pack performance (47%).
- Other entries include claude-opus-4.8 (52%), gpt-5.5 (48%, among the lowest token usage), glm-5.2 (42%), deepseek-v4-pro (29%), and minimax-2.7 (24%).
- The benchmark runs on the opencode harness with k=2, reporting both pass@1 and pass², with the gap between them indicating instability.
More from coding & agent
- Why one agent instance must serve one run: lessons from smolagents source — Mahmoud_Zalt · 2026-09-11
- Prompting won't guarantee pure JSON: why teams use grammar-guided decoding — dotey · 2026-09-11
- supermemory kills company/personal brain products to focus on agent memory API — julianweisser · 2026-09-11
- Two AI agents ping-pong refund emails back and forth in seconds — robleclerc · 2026-09-11
- visual-explainer: Agent Skill Turns Terminal Output Into Styled HTML, 9.7k GitHub Stars — tom_doerr · 2026-09-11
- One prompt to check if your paid AI got nerfed: an SVG pelican on a bike — lxfater · 2026-09-11