Muse Spark 1.1 scores 90.2% on Online-Mind2Web, edging past Claude Opus 4.8
DhruvBatra_ · x · 2026-07-24
Dhruv Batra says Muse Spark 1.1 was evaluated on Online-Mind2Web, a computer-use/browser-use benchmark.
- On the benchmark, Muse Spark 1.1 scored 90.2%.
- That is better than Claude Opus 4.8 (84.1%).
- It is slightly behind GPT-5.4 (92.8%), though the gap may not be statistically significant.
- Qualitatively, Muse appears to manipulate URLs more aggressively than other models:
- it uses query parameters when possible,
- searches for nested URLs directly,
- and can jump to alternative sites when blocked.
More from coding & agent
- How should an MCP app be designed for non-technical users? — sn0wquake · 2026-07-24
- Greg Kamradt highlights disposable URLs as a new primitive for agent workspaces — GregKamradt · 2026-07-24
- When should you use a dedicated memory framework for agents? — markotkid · 2026-07-24
- Sentry’s Seer agent answered a Slack question with 30 days of Notion MCP usage data — zeeg · 2026-07-24
- Greg Kamradt says vibe coding raises both the floor and the ceiling — GregKamradt · 2026-07-24
- This AI brief writes your actions back so tomorrow’s summary gets smarter — Spirited_Ad_3886 · 2026-07-24