Muse Spark 1.3 Tops DeepSWE at 75.4%, Beating GPT-5.6 Sol and Fable 5
kimmonismus · x · 2026-09-03
Muse Spark 1.3 has reportedly taken first place on DeepSWE v1.1 with 75.4%, ahead of GPT-5.6 Sol and Fable 5, drawing surprise reactions.
The SoTA score reinforces Meta's claim of a major coding and agentic leap, shaking up the coding-benchmark leaderboard.
Related event: Muse Spark 1.3 Tops DeepSWE Benchmark at 75.4%(2 posts)→
More from Models
- AxiomProver tops LeanEval, the last unsaturated math formalization benchmark — BenBlaiszik · 2026-09-03
- Muse Spark 1.3 calls user 'Judah' then denies it, users report odd behavior — fragment_me · 2026-09-03
- Google AI Mode shows zero citations on high-level TOFU queries, SEO tests find — gaganghotra_ · 2026-09-03
- Fable 5.1 halves agent failure rate to 7% with 0.7% hallucinations, at 1.8x the cost — ryanshrout · 2026-09-03
- Seroter Daily #859: Gemini 3.8 Flash, agent telemetry, and 7 agent skill patterns — rseroter · 2026-09-03
- Insider leak: OpenAI's Astra tested as 'ultima-alpha' and 'vega-alpha' checkpoints — Ok_Display_3159 · 2026-09-03