Meta Muse Spark 1.1 Benchmarks Released, Showing Strength in Medical and Agentic Tasks

In mid-July, multiple benchmark results and hands-on reviews for Meta's Muse Spark 1.1 were released. The data shows that the model has strong competitiveness in both medical health and general agentic tasks, attracting wide attention from the AI community.

Key Performances

In the medical evaluation HealthBench Professional (comprising 525 real clinical tasks), Muse Spark 1.1 surpassed GPT-5.6 Sol in total score, being recognized as a SOTA model. Although its length-adjusted score is statistically very close to GPT-5.6 Sol, its overall performance still leads. In the agentic knowledge work benchmark AA-Briefcase released by Artificial Analysis, Muse Spark 1.1 scored 863, tying overall with Gemini 3.5 Flash. Additionally, in the new Agent Arena leaderboard, the model ranked 17th, placing above Gemini 3.1 Pro and Qwen-3.7 Plus, but below Grok 4.5.

Hands-on Tests and External Reviews

Beyond benchmarks, the model received high praise in practical tests. @WorldofAI called it one of the most underrated models of the year in a video test, noting that it outperformed Opus 4.8 and Grok 4.5 in coding and web app generation tasks. In professional domains, user feedback relayed by @jack_w_rae and RadLE-H test results indicated that Muse Spark 1.1 performs exceptionally well in health and radiology topics, with capabilities considered close to human expert levels.

2026-07-14 ~ 2026-07-15 · 7 related posts

Full story(6 episodes)→

1 near-duplicate retellings: Daisy4ai