Meta Muse Spark 1.1 Shines Across Multiple Benchmarks

Meta's newly released Muse Spark 1.1 has demonstrated robust capabilities across multiple benchmarks and practical applications, attracting widespread attention for its extreme cost-efficiency and performance that closely approaches frontier models.

Benchmark Performance and Core Capabilities

Muse Spark 1.1 achieved notable breakthroughs in both text and coding skills. According to Arena data, the model scored 1494 to rank 5th in Text Arena, up 7 points from the original version, successfully entering the Pareto frontier. It improved in 12 out of 15 selected categories, with its Expert ranking jumping from #36 to #15, alongside significant gains in instruction following. In Code Arena's frontend category, it ranked 9th overall and reached 2nd place in Data & Analytics. Furthermore, data shared by @shuchaobi indicates a massive leap in formal math capabilities, with its ProofBench score surging from 17% to 39%. It also ranked 4th on the Vals Index leaderboard while operating as the fastest model with the lowest latency in the top ten.

Cost-Efficiency and Industry Feedback

The extreme value for money of Muse Spark 1.1 has become a major talking point. @alexandr Wang noted that its output token cost is about 90% cheaper than Fable, with a comprehensive cost of around $3.5/M. Evaluations from Artificial Analysis show that Muse Spark 1.1 (xhigh) scored 69 on the Coding Agent Index, demonstrating excellent cost-efficiency under the Opencode harness. In terms of hands-on experience, user feedback highlights strong frontend design capabilities and solid agentic skills. Given the same prompts, its output is considered more refined and interactive, with overall performance approaching ChatGPT Sol Medium, Claude Opus 4.8, and Grok 4.5. Some suggest that with both Grok 4.5 and Muse performing near top-tier levels, Anthropic would be at a disadvantage if they removed Fable from subscriptions. Additionally, a benchmark summary claimed its performance on SciCode surpasses GPT-5.6 with a score of 58%.

2026-07-11 ~ 2026-07-12 · 13 related posts

Full story(6 episodes)→