FULL STORY

Meta Muse Spark 1.1: From Teaser to Real-World Tests

Meta released the multimodal model Muse Spark 1.1, leading with highly competitive pricing and strong agent capabilities. Real-world tests highlight its exceptional value and top performance in vertical domains like healthcare, despite trailing overall leaders.

2026-07-03 ~ 2026-07-15 · 6 episodes · 172 posts

Episode 1 · Meta Superintelligence Labs: Muse Spark Update and High-Performance Opus-Level Variant Coming Soon (2026-07-03, 2 posts)

Episode 2 · Meta Launches Muse Spark 1.1: A Low-Cost, High-Performance Agentic Model (2026-07-09, 133 posts)

On July 9, Meta officially launched Muse Spark 1.1 (formerly Hornbill), a multimodal reasoning model. Upgraded from its first version, it focuses on strong agentic and coding capabilities at a low price, signaling Meta's return to the frontier AI race. The model is now accessible to developers via the public beta Meta Model API and Meta AI.

Core Capabilities and Technical Details

Muse Spark 1.1 features a 1 million token context and can compress history to retain essential steps. Meta highlighted improvements in tool use, computer use, coding, and multimodal understanding. It can zero-shot generalize to new tools, orchestrate multi-agent tasks, maintain context across apps, and process visual and audio inputs.

Benchmark Performance and Reactions

According to Alexandr Wang, the model achieved SOTA on benchmarks like Harvey Legal Bench and TaxEval. Its agentic capabilities rival GPT-5.5 and Opus-4.8, even outperforming Opus 4.8 and Grok 4.5 in some out-of-distribution evaluations. Both Garry Tan and Wang reported excellent real-world results using it on OpenClaw (robotics scenarios). Industry figures like Bindu Reddy also praised its benchmark performance and API availability.

Internal Use and Ecosystem Support

Meta stated that Muse Spark 1.1 is already used for internal coding and research workflows, automating model development. Externally, early adopters including Replit and Box have started building on the new API.

113 more related posts →

Episode 3 · Meta Muse Spark 1.1 Shines in Benchmarks with Major Gains in Intelligence and Coding (2026-07-09, 15 posts)

Meta's native multimodal reasoning model, Muse Spark 1.1, has garnered significant attention following comprehensive benchmarks by Artificial Analysis. The evaluation reveals substantial improvements in core capabilities, positioning the model as highly competitive in both performance and cost efficiency.

Key Benchmarks and Capability Improvements

According to Artificial Analysis, Muse Spark 1.1 scores 51 on the Intelligence Index, an 8-point increase over version 1.0. This growth is primarily driven by enhancements in scientific reasoning, coding, and knowledge. In specific benchmarks, it achieved 58% on SciCode (ranking 3rd, behind Claude Fable 5 and Gemini 3.1) and 92.9% on CyBench (closely trailing Claude Opus 4.6's 93%). Additionally, it scored 69 on the Coding Agent Index using the Opencode harness. Notably, the model's intelligence ranking improvement is largely attributed to a reduction in hallucination rates, contrasting with Grok 4.5's pattern of simultaneous increases in both accuracy and hallucinations.

Computer Use and Practical Feedback

Muse Spark 1.1 is evaluated as "very strong" in computer use (CUA). Practical feedback highlights its proficiency in front-end generation and overall output quality, bringing pleasant surprises for "vibe-coding." However, users also noted minor shortcomings, such as failing to check if a file already exists before overwriting it. Overall, the community considers it more than sufficient for many tasks, coupled with extremely fast speeds.

Cost and Efficiency Advantages

Beyond raw capability, Muse Spark 1.1 demonstrates exceptional token efficiency and cost control. Running the entire Intelligence Index consumed only 94 million output tokens, significantly lower than comparable models. Its operational efficiency is regarded as superior to Kimi K2.6 or GLM 5.2, effectively filling a market gap between high performance and low cost.

Episode 4 · Meta Muse Spark 1.1 Shines Across Multiple Benchmarks (2026-07-11, 13 posts)

Meta's newly released Muse Spark 1.1 has demonstrated robust capabilities across multiple benchmarks and practical applications, attracting widespread attention for its extreme cost-efficiency and performance that closely approaches frontier models.

Benchmark Performance and Core Capabilities

Muse Spark 1.1 achieved notable breakthroughs in both text and coding skills. According to Arena data, the model scored 1494 to rank 5th in Text Arena, up 7 points from the original version, successfully entering the Pareto frontier. It improved in 12 out of 15 selected categories, with its Expert ranking jumping from #36 to #15, alongside significant gains in instruction following. In Code Arena's frontend category, it ranked 9th overall and reached 2nd place in Data & Analytics. Furthermore, data shared by @shuchaobi indicates a massive leap in formal math capabilities, with its ProofBench score surging from 17% to 39%. It also ranked 4th on the Vals Index leaderboard while operating as the fastest model with the lowest latency in the top ten.

Cost-Efficiency and Industry Feedback

The extreme value for money of Muse Spark 1.1 has become a major talking point. @alexandr Wang noted that its output token cost is about 90% cheaper than Fable, with a comprehensive cost of around $3.5/M. Evaluations from Artificial Analysis show that Muse Spark 1.1 (xhigh) scored 69 on the Coding Agent Index, demonstrating excellent cost-efficiency under the Opencode harness. In terms of hands-on experience, user feedback highlights strong frontend design capabilities and solid agentic skills. Given the same prompts, its output is considered more refined and interactive, with overall performance approaching ChatGPT Sol Medium, Claude Opus 4.8, and Grok 4.5. Some suggest that with both Grok 4.5 and Muse performing near top-tier levels, Anthropic would be at a disadvantage if they removed Fable from subscriptions. Additionally, a benchmark summary claimed its performance on SciCode surpasses GPT-5.6 with a score of 58%.

Episode 5 · Meta Muse Spark 1.1 Benchmarks Released, Showing Strength in Medical and Agentic Tasks (2026-07-14, 7 posts)

In mid-July, multiple benchmark results and hands-on reviews for Meta's Muse Spark 1.1 were released. The data shows that the model has strong competitiveness in both medical health and general agentic tasks, attracting wide attention from the AI community.

Key Performances

In the medical evaluation HealthBench Professional (comprising 525 real clinical tasks), Muse Spark 1.1 surpassed GPT-5.6 Sol in total score, being recognized as a SOTA model. Although its length-adjusted score is statistically very close to GPT-5.6 Sol, its overall performance still leads. In the agentic knowledge work benchmark AA-Briefcase released by Artificial Analysis, Muse Spark 1.1 scored 863, tying overall with Gemini 3.5 Flash. Additionally, in the new Agent Arena leaderboard, the model ranked 17th, placing above Gemini 3.1 Pro and Qwen-3.7 Plus, but below Grok 4.5.

Hands-on Tests and External Reviews

Beyond benchmarks, the model received high praise in practical tests. @WorldofAI called it one of the most underrated models of the year in a video test, noting that it outperformed Opus 4.8 and Grok 4.5 in coding and web app generation tasks. In professional domains, user feedback relayed by @jack_w_rae and RadLE-H test results indicated that Muse Spark 1.1 performs exceptionally well in health and radiology topics, with capabilities considered close to human expert levels.

Episode 6 · Meta Launches Muse Spark 1.1 with API Access (2026-07-15, 2 posts)

Meta officially launched Muse Spark 1.1, a major upgrade designed for agentic coding and computer-use workflows, accompanied by newly available API access.