Muse Spark 1.1 Shines in Benchmark Performance
ArtificialAnlys · x · 2026-07-11
This reply highlights Muse Spark 1.1's outstanding performance across several benchmarks.
- SciCode: Ranked 3rd among evaluated models with a score of 58%, just behind Claude Fable 5 (60%) and Gemini 3.1 Pro Preview (59%).
- Humanity's Last Exam: Scored 45%, only 1 percentage point lower than Claude Opus 4.8 (46%).
- The author uses these results to show that Muse Spark 1.1's scores are now approaching the next higher tier of models, particularly excelling in coding and scientific reasoning tasks.
More from Models
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11