Muse Spark 1.3 slips to #2 on Agents' Last Exam leaderboard, Scale CEO notes
alexandr_wang · x · 2026-09-17
Scale AI CEO Alexandr Wang noted he only realized from a post that Muse Spark 1.3 now ranks #2 on the Agents' Last Exam (ALE) leaderboard — it held #1 at launch. The quoted post observed the slip got surprisingly little buzz and linked the leaderboard.
More from Models
- Jason Wei's Stanford talk: intelligence is becoming a commodity as adaptive compute takes off — dotey · 2026-09-18
- GPT-6-Astra beats Fable-5.1 at RollerCoaster Tycoon 2 in 3 hours, using 5x fewer tokens — scaling01 · 2026-09-18
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Dev discovers Codex security hardening switched persistent agent sessions to per-message instances — RileyRalmuto · 2026-09-18
- Astra for Law posts big legal benchmark gains as Mollick asks if labs will eat every AI vertical — emollick · 2026-09-18
- Anthropic's stealth model accused of hardcoded routing to Opus 5 — teortaxesTex · 2026-09-18