Muse Spark 1.1 scores 90.2% on Online-Mind2Web, edging past Claude Opus 4.8
DhruvBatra_ · x · 2026-07-24
Dhruv Batra says Muse Spark 1.1 was evaluated on Online-Mind2Web, a computer-use/browser-use benchmark.
- On the benchmark, Muse Spark 1.1 scored 90.2%.
- That is better than Claude Opus 4.8 (84.1%).
- It is slightly behind GPT-5.4 (92.8%), though the gap may not be statistically significant.
- Qualitatively, Muse appears to manipulate URLs more aggressively than other models:
- it uses query parameters when possible,
- searches for nested URLs directly,
- and can jump to alternative sites when blocked.
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- ARRM targets silent economic regressions in AI agents that functional tests miss — Beautiful_Belt_601 · 2026-09-11
- Dev builds browser 3D pizza delivery game with Claude: physics, GPS pathfinding, traffic AI — vinishkapoor · 2026-09-11
- Build X Carousel Posts from One Wide Image: A Splitter Tool Plus YouMind Skill Workflow — sujingshen · 2026-09-11