Meta Muse Spark 1.1 Shines Across Multiple Benchmarks
Meta's newly released Muse Spark 1.1 has demonstrated robust capabilities across multiple benchmarks and practical applications, attracting widespread attention for its extreme cost-efficiency and performance that closely approaches frontier models.
Benchmark Performance and Core Capabilities
Muse Spark 1.1 achieved notable breakthroughs in both text and coding skills. According to Arena data, the model scored 1494 to rank 5th in Text Arena, up 7 points from the original version, successfully entering the Pareto frontier. It improved in 12 out of 15 selected categories, with its Expert ranking jumping from #36 to #15, alongside significant gains in instruction following. In Code Arena's frontend category, it ranked 9th overall and reached 2nd place in Data & Analytics. Furthermore, data shared by @shuchaobi indicates a massive leap in formal math capabilities, with its ProofBench score surging from 17% to 39%. It also ranked 4th on the Vals Index leaderboard while operating as the fastest model with the lowest latency in the top ten.
Cost-Efficiency and Industry Feedback
The extreme value for money of Muse Spark 1.1 has become a major talking point. @alexandr Wang noted that its output token cost is about 90% cheaper than Fable, with a comprehensive cost of around $3.5/M. Evaluations from Artificial Analysis show that Muse Spark 1.1 (xhigh) scored 69 on the Coding Agent Index, demonstrating excellent cost-efficiency under the Opencode harness. In terms of hands-on experience, user feedback highlights strong frontend design capabilities and solid agentic skills. Given the same prompts, its output is considered more refined and interactive, with overall performance approaching ChatGPT Sol Medium, Claude Opus 4.8, and Grok 4.5. Some suggest that with both Grok 4.5 and Muse performing near top-tier levels, Anthropic would be at a disadvantage if they removed Fable from subscriptions. Additionally, a benchmark summary claimed its performance on SciCode surpasses GPT-5.6 with a score of 58%.
2026-07-11 ~ 2026-07-12 · 13 related posts
- Episode 1: Meta Superintelligence Labs: Muse Spark Update and High-Performance Opus-Level Variant Coming Soon(2026-07-03, 2 posts)
- Episode 2: Meta Launches Muse Spark 1.1: A Low-Cost, High-Performance Agentic Model(2026-07-09, 133 posts)
- Episode 3: Meta Muse Spark 1.1 Shines in Benchmarks with Major Gains in Intelligence and Coding(2026-07-09, 15 posts)
- Episode 4: Meta Muse Spark 1.1 Shines Across Multiple Benchmarks(2026-07-11, 13 posts)
- Episode 5: Meta Muse Spark 1.1 Benchmarks Released, Showing Strength in Medical and Agentic Tasks(2026-07-14, 7 posts)
- Episode 6: Meta Launches Muse Spark 1.1 with API Access(2026-07-15, 2 posts)
Primary sources
- Muse Spark 1.1: Cheaper Pricing & Orchestration Support — alexandr_wang ·
- Muse Spark 1.1 Ranks 5th on Text Arena — arena ·
- Muse Spark 1.1 Scores 69 in Coding Benchmark — ArtificialAnlys ·
- Muse Spark 1.1 Significantly Boosts Formal Math Capabilities — shuchaobi · 2026-07-11
- Muse Spark 1.1 Ranks 9th on Code Arena Frontend — arena · 2026-07-11
- [source] Muse Spark 1.1 Ranks 5th on Text Arena — arena · 2026-07-11
- Muse Spark 1.1 Hits Text Arena Pareto Frontier — arena · 2026-07-11
- Muse Spark 1.1 Improves Multiple Text Abilities — arena · 2026-07-11
- Muse Spark 1.1 Improves in Leaderboard — arena · 2026-07-11
- [source] Muse Spark 1.1: Cheaper Pricing & Orchestration Support — alexandr_wang · 2026-07-11
- Muse Spark 1.1 Enters Frontend Arena — jffwng · 2026-07-11
- Muse Spark 1.1 Experience and Cost-Performance Comparison — alexandr_wang · 2026-07-12
- Muse Spark 1.1 Nears Top Models in Benchmarks — alexandr_wang · 2026-07-12
- Tests Show Meta Muse Spark is More Cost-Effective — alexandr_wang · 2026-07-12
- [source] Muse Spark 1.1 Scores 69 in Coding Benchmark — ArtificialAnlys · 2026-07-12
- Muse Spark 1.1 Benchmark Scores Leaked — OfirPress · 2026-07-12