Muse Spark 1.1 IQ Ranking Rises by Reducing Hallucinations

ArtificialAnlys · x · 2026-07-11

Artificial Analysis observed that Muse Spark 1.1's ranking improvement pattern on the Intelligence Index is the opposite of Grok 4.5. Grok 4.5's improvement came with a simultaneous rise in both accuracy and hallucination rates, whereas Muse Spark 1.1's jump from 4th to 18th place is primarily due to an "abstention" mechanism: with accuracy remaining basically flat, its hallucination rate dropped significantly by 35 percentage points to 38%.

Additionally, when running the Intelligence Index tests, Muse Spark 1.1 demonstrated better token efficiency, consuming 94 million output tokens, which is lower than similarly-scoring models like GPT-5.4 (xhigh) and GLM-5.2 (max).

Related event: Meta Muse Spark 1.1 Shines in Benchmarks with Major Gains in Intelligence and Coding(15 posts)→

Original post →

More from Models

Models channel →