Muse Spark 1.1 Ranks 17th in Benchmarks

arena · x · 2026-07-15

Agent Arena has released a new leaderboard and evaluation guide. Meta's Muse Spark 1.1 ranks 17th, placing it above Gemini 3.1 Pro and Qwen-3.7 Plus, but below Grok 4.5 and GLM 5.2.

The rankings are based on Agent Arena, which leverages millions of real, long-horizon agent tasks submitted by global users to measure model performance. Models are allowed to use tools like web search, file systems, and terminals to complete complex workflows. Rankings are determined by task outcomes relative to an average model, measured using causal tracing methods.

Related event: Meta Muse Spark 1.1 Benchmarks Released, Showing Strength in Medical and Agentic Tasks(7 posts)→

Original post →

More from Models

Models channel →