Muse Spark 1.1 Shines in Benchmark Performance
ArtificialAnlys · x · 2026-07-11
This reply highlights Muse Spark 1.1's outstanding performance across several benchmarks.
- SciCode: Ranked 3rd among evaluated models with a score of 58%, just behind Claude Fable 5 (60%) and Gemini 3.1 Pro Preview (59%).
- Humanity's Last Exam: Scored 45%, only 1 percentage point lower than Claude Opus 4.8 (46%).
- The author uses these results to show that Muse Spark 1.1's scores are now approaching the next higher tier of models, particularly excelling in coding and scientific reasoning tasks.
More from Models
- Kimi K3 rises to No. 4 on the Agent Arena leaderboard — HeyZoyaKhan · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Google DeepMind launches Gemini 3.5 Flash Cyber for faster, cheaper code security — ralucaadapopa · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22