Meta Muse Spark 1.1 Shines Across Multiple Benchmarks
Meta's newly released Muse Spark 1.1 has demonstrated robust capabilities across multiple benchmarks and practical applications, attracting widespread attention for its extreme cost-efficiency and performance that closely approaches frontier models.
Benchmark Performance and Core Capabilities
Muse Spark 1.1 achieved notable breakthroughs in both text and coding skills. According to Arena data, the model scored 1494 to rank 5th in Text Arena, up 7 points from the original version, successfully entering the Pareto frontier. It improved in 12 out of 15 selected categories, with its Expert ranking jumping from #36 to #15, alongside significant gains in instruction following. In Code Arena's frontend category, it ranked 9th overall and reached 2nd place in Data & Analytics. Furthermore, data shared by @shuchaobi indicates a massive leap in formal math capabilities, with its ProofBench score surging from 17% to 39%. It also ranked 4th on the Vals Index leaderboard while operating as the fastest model with the lowest latency in the top ten.
Cost-Efficiency and Industry Feedback
The extreme value for money of Muse Spark 1.1 has become a major talking point. @alexandr Wang noted that its output token cost is about 90% cheaper than Fable, with a comprehensive cost of around $3.5/M. Evaluations from Artificial Analysis show that Muse Spark 1.1 (xhigh) scored 69 on the Coding Agent Index, demonstrating excellent cost-efficiency under the Opencode harness. In terms of hands-on experience, user feedback highlights strong frontend design capabilities and solid agentic skills. Given the same prompts, its output is considered more refined and interactive, with overall performance approaching ChatGPT Sol Medium, Claude Opus 4.8, and Grok 4.5. Some suggest that with both Grok 4.5 and Muse performing near top-tier levels, Anthropic would be at a disadvantage if they removed Fable from subscriptions. Additionally, a benchmark summary claimed its performance on SciCode surpasses GPT-5.6 with a score of 58%.
2026-07-11 ~ 2026-07-12 · 13 related posts
- Episode 1: Meta超级智能实验室:Muse Spark更新及Opus级高性能变体即将发布(2026-07-03, 2 posts)
- Episode 2: Meta发布Muse Spark 1.1,主打低成本与强Agentic能力(2026-07-09, 133 posts)
- Episode 3: Meta Muse Spark 1.1多基准评测亮眼,智能与编码能力大幅提升(2026-07-09, 15 posts)
- Episode 4: Meta Muse Spark 1.1 多项榜单成绩亮眼(2026-07-11, 13 posts)
- Episode 5: Meta Muse Spark 1.1 多项评测出炉,医疗与 Agent 表现亮眼(2026-07-14, 7 posts)
- Episode 6: Meta发布Muse Spark 1.1并开放API(2026-07-15, 2 posts)
- Muse Spark 1.1 Significantly Boosts Formal Math Capabilities — shuchaobi · 2026-07-11
- Muse Spark 1.1 Ranks 9th on Code Arena Frontend — arena · 2026-07-11
- [source] Muse Spark 1.1 Ranks 5th on Text Arena — arena · 2026-07-11
- Muse Spark 1.1 Hits Text Arena Pareto Frontier — arena · 2026-07-11
- Muse Spark 1.1 Improves Multiple Text Abilities — arena · 2026-07-11
- Muse Spark 1.1 Improves in Leaderboard — arena · 2026-07-11
- [source] Muse Spark 1.1: Cheaper Pricing & Orchestration Support — alexandr_wang · 2026-07-11
- Muse Spark 1.1 Enters Frontend Arena — jffwng · 2026-07-11
- Muse Spark 1.1 Experience and Cost-Performance Comparison — alexandr_wang · 2026-07-12
- Muse Spark 1.1 Nears Top Models in Benchmarks — alexandr_wang · 2026-07-12
- Tests Show Meta Muse Spark is More Cost-Effective — alexandr_wang · 2026-07-12
- [source] Muse Spark 1.1 Scores 69 in Coding Benchmark — ArtificialAnlys · 2026-07-12
- Muse Spark 1.1 Benchmark Scores Leaked — OfirPress · 2026-07-12