Muse Spark 1.1 Ranks 17th in Benchmarks
arena · x · 2026-07-15
Agent Arena has released a new leaderboard and evaluation guide. Meta's Muse Spark 1.1 ranks 17th, placing it above Gemini 3.1 Pro and Qwen-3.7 Plus, but below Grok 4.5 and GLM 5.2.
The rankings are based on Agent Arena, which leverages millions of real, long-horizon agent tasks submitted by global users to measure model performance. Models are allowed to use tools like web search, file systems, and terminals to complete complex workflows. Rankings are determined by task outcomes relative to an average model, measured using causal tracing methods.
More from Models
- Open-Weight Model Hy3 Ranks #16 in Frontend Code Arena — arena · 2026-07-22
- DeepSWE Eval: Kimi K3 Matches Claude Fable 5 at 35% of the Cost — togethercompute · 2026-07-22
- Gemini 3.5 Flash Outperforms GPT-5.6 in Light Coding Tasks — Shick_hydro · 2026-07-22
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- Gemma-4-26B-a4B reportedly beats Qwen3.6 and Qwen3.5 MoE fine-tunes — JLeonsarmiento · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22