Muse Spark 1.1 IQ Ranking Rises by Reducing Hallucinations
ArtificialAnlys · x · 2026-07-11
Artificial Analysis observed that Muse Spark 1.1's ranking improvement pattern on the Intelligence Index is the opposite of Grok 4.5. Grok 4.5's improvement came with a simultaneous rise in both accuracy and hallucination rates, whereas Muse Spark 1.1's jump from 4th to 18th place is primarily due to an "abstention" mechanism: with accuracy remaining basically flat, its hallucination rate dropped significantly by 35 percentage points to 38%.
Additionally, when running the Intelligence Index tests, Muse Spark 1.1 demonstrated better token efficiency, consuming 94 million output tokens, which is lower than similarly-scoring models like GPT-5.4 (xhigh) and GLM-5.2 (max).
More from Models
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- OpenAI’s Codex + GPT-5.6 Sol hits 99% recall in Project APE verification tests — soumitrashukla9 · 2026-07-22
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22