Alexandr Wang flags Muse Spark 1.3 eval: time horizon now matches GPT-5.6 Sol and Opus 5
alexandr_wang · x · 2026-09-07
Scale AI founder Alexandr Wang boosted aquilesfd's eval of Muse Spark 1.3 max, calling it an interesting one.
Key claims from the quoted post: Muse Spark 1.3 max is a respectable upgrade over xhigh, improving performance across many categories at slightly higher cost and latency per task. Its time horizon improved massively, now on par with GPT-5.6 Sol and Opus 5.
More from Models
- Astra's AGI estimate jumps with tool use — is 'ASI already here' just a harness question? — kevinnbass · 2026-09-07
- GLM 5.3 and Qwen 3.8 now run really well locally on single desktops — jasonkneen · 2026-09-07
- New benchmark probes LLM self-modeling: RL lifts open models but counterfactual errors persist — dair_ai · 2026-09-07
- Cool presentation aside, Astra still can't nail research-level single-step reasoning — xiaosun86 · 2026-09-07
- 31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day — ionutvi · 2026-09-07
- GPT-6 Astra generates missile evasion simulation, showcasing stunning capability — algo_diver · 2026-09-07