Meta Muse Spark 1.1 Shines in Benchmarks with Major Gains in Intelligence and Coding

Meta's native multimodal reasoning model, Muse Spark 1.1, has garnered significant attention following comprehensive benchmarks by Artificial Analysis. The evaluation reveals substantial improvements in core capabilities, positioning the model as highly competitive in both performance and cost efficiency.

Key Benchmarks and Capability Improvements

According to Artificial Analysis, Muse Spark 1.1 scores 51 on the Intelligence Index, an 8-point increase over version 1.0. This growth is primarily driven by enhancements in scientific reasoning, coding, and knowledge. In specific benchmarks, it achieved 58% on SciCode (ranking 3rd, behind Claude Fable 5 and Gemini 3.1) and 92.9% on CyBench (closely trailing Claude Opus 4.6's 93%). Additionally, it scored 69 on the Coding Agent Index using the Opencode harness. Notably, the model's intelligence ranking improvement is largely attributed to a reduction in hallucination rates, contrasting with Grok 4.5's pattern of simultaneous increases in both accuracy and hallucinations.

Computer Use and Practical Feedback

Muse Spark 1.1 is evaluated as "very strong" in computer use (CUA). Practical feedback highlights its proficiency in front-end generation and overall output quality, bringing pleasant surprises for "vibe-coding." However, users also noted minor shortcomings, such as failing to check if a file already exists before overwriting it. Overall, the community considers it more than sufficient for many tasks, coupled with extremely fast speeds.

Cost and Efficiency Advantages

Beyond raw capability, Muse Spark 1.1 demonstrates exceptional token efficiency and cost control. Running the entire Intelligence Index consumed only 94 million output tokens, significantly lower than comparable models. Its operational efficiency is regarded as superior to Kimi K2.6 or GLM 5.2, effectively filling a market gap between high performance and low cost.

2026-07-09 ~ 2026-07-11 · 15 related posts

Full story(6 episodes)→

1 near-duplicate retellings: shuchaobi