Muse Spark 1.3 tops DeepSWE at 75.4%, beating GPT-5.6 Sol and Fable 5
kimmonismus · x · 2026-09-03
Blogger kimmonismus notes that the newly released Muse Spark 1.3 takes first place on the DeepSWE coding benchmark with 75.4%, ahead of GPT-5.6 Sol and Fable 5. For reference, Fable 5 sits at 70%, and there appear to be no published evals yet for Fable 5.1.
Related event: Muse Spark 1.3 Tops DeepSWE Benchmark at 75.4%(2 posts)→
More from Models
- Counterfactual: without reasoning models, AI today might just be reaching o3-level — Jsevillamol · 2026-09-03
- Anthropic's Fable 5.1 hits 90% on ARC-AGI-2 at 32% lower cost per task than Fable 5 — rohanpaul_ai · 2026-09-03
- Team shares 4 real LLM uses: contract negotiation, agent clarification, grading, math — xuanalogue · 2026-09-03
- Team claims h3 max is the undisputed #1 frontier video model across benchmarks — isidentical · 2026-09-03
- Meta's SAM 3, with image and video segmentation, tops Hugging Face trending — facebook · 2026-09-03
- Anthropic internal 'retirement home' for old Claude models sparks confabulation concerns — repligate · 2026-09-03