Meta Muse Spark 1.1 Shines in Benchmarks with Major Gains in Intelligence and Coding
Meta's native multimodal reasoning model, Muse Spark 1.1, has garnered significant attention following comprehensive benchmarks by Artificial Analysis. The evaluation reveals substantial improvements in core capabilities, positioning the model as highly competitive in both performance and cost efficiency.
Key Benchmarks and Capability Improvements
According to Artificial Analysis, Muse Spark 1.1 scores 51 on the Intelligence Index, an 8-point increase over version 1.0. This growth is primarily driven by enhancements in scientific reasoning, coding, and knowledge. In specific benchmarks, it achieved 58% on SciCode (ranking 3rd, behind Claude Fable 5 and Gemini 3.1) and 92.9% on CyBench (closely trailing Claude Opus 4.6's 93%). Additionally, it scored 69 on the Coding Agent Index using the Opencode harness. Notably, the model's intelligence ranking improvement is largely attributed to a reduction in hallucination rates, contrasting with Grok 4.5's pattern of simultaneous increases in both accuracy and hallucinations.
Computer Use and Practical Feedback
Muse Spark 1.1 is evaluated as "very strong" in computer use (CUA). Practical feedback highlights its proficiency in front-end generation and overall output quality, bringing pleasant surprises for "vibe-coding." However, users also noted minor shortcomings, such as failing to check if a file already exists before overwriting it. Overall, the community considers it more than sufficient for many tasks, coupled with extremely fast speeds.
Cost and Efficiency Advantages
Beyond raw capability, Muse Spark 1.1 demonstrates exceptional token efficiency and cost control. Running the entire Intelligence Index consumed only 94 million output tokens, significantly lower than comparable models. Its operational efficiency is regarded as superior to Kimi K2.6 or GLM 5.2, effectively filling a market gap between high performance and low cost.
2026-07-09 ~ 2026-07-11 · 15 related posts
- Episode 1: Meta超级智能实验室:Muse Spark更新及Opus级高性能变体即将发布(2026-07-03, 2 posts)
- Episode 2: Meta发布Muse Spark 1.1,主打低成本与强Agentic能力(2026-07-09, 133 posts)
- Episode 3: Meta Muse Spark 1.1多基准评测亮眼,智能与编码能力大幅提升(2026-07-09, 15 posts)
- Episode 4: Meta Muse Spark 1.1 多项榜单成绩亮眼(2026-07-11, 13 posts)
- Episode 5: Meta Muse Spark 1.1 多项评测出炉,医疗与 Agent 表现亮眼(2026-07-14, 7 posts)
- Episode 6: Meta发布Muse Spark 1.1并开放API(2026-07-15, 2 posts)
- Muse Spark 1.1 Scores Near the Top — scaling01 · 2026-07-09
- Muse Spark 1.1 Coding Capabilities Win Praise — alexandr_wang · 2026-07-10
- Muse Spark 1.1 Receives Positive Early Feedback — BenBajarin · 2026-07-10
- [source] Meta Muse Spark 1.1 Evaluation Released — ArtificialAnlys · 2026-07-11
- Muse Spark 1.1 Evaluation Score Jumps 8 Points — ArtificialAnlys · 2026-07-11
- Muse Spark 1.1 Shines in Benchmark Performance — ArtificialAnlys · 2026-07-11
- Muse Spark 1.1 Comes at a Lower Cost — ArtificialAnlys · 2026-07-11
- Muse Spark 1.1 IQ Ranking Rises by Reducing Hallucinations — ArtificialAnlys · 2026-07-11
- Muse Spark 1.1 Cuts Hallucination Rate — ArtificialAnlys · 2026-07-11
- Muse Spark 1.1 Evaluation Results — elemental-mind · 2026-07-11
- Muse Spark 1.1 Shows Significant Coding Improvements — EdwardSun0909 · 2026-07-11
- [source] Muse Spark 1.1 Computer Use Enhanced — alexandr_wang · 2026-07-11
- [source] Muse Spark 1.1 Ranks High in Coding Agent Eval — ArtificialAnlys · 2026-07-11
- Meta Muse Spark 1.1 Hands-on Review — alexandr_wang · 2026-07-11
1 near-duplicate retellings: shuchaobi