Muse Spark 1.1 Debuts at 6th on APEX-Agents
giffmana · x · 2026-07-14
Muse Spark 1.1 has made its debut on APEX-Agents, ranking 6th.
- Pass@1: 37.1%, Mean criteria passed: 52.0%.
- Officials noted that in about 10% of the tasks, the model failed before completing the trajectory, resulting in an automatic score of 0 for these tasks.
- If trajectories impacted by Meta's aggressive content filtering are excluded, Pass@1 would rise to 41.4%, potentially ranking it ahead of GPT-5.6 Sol.
- This benchmark is designed for long-horizon tasks based on real-world banking, legal, and consulting work.
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- giffmana: the env being used in training is part of the point — giffmana · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11