Muse Spark 1.1 Ranks 17th in Benchmarks
arena · x · 2026-07-15
Agent Arena has released a new leaderboard and evaluation guide. Meta's Muse Spark 1.1 ranks 17th, placing it above Gemini 3.1 Pro and Qwen-3.7 Plus, but below Grok 4.5 and GLM 5.2.
The rankings are based on Agent Arena, which leverages millions of real, long-horizon agent tasks submitted by global users to measure model performance. Models are allowed to use tools like web search, file systems, and terminals to complete complex workflows. Rankings are determined by task outcomes relative to an average model, measured using causal tracing methods.
More from Models
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11