Muse Spark 1.1 Climbs Debate Benchmark Rankings

alexandr_wang · x · 2026-07-14

Muse Spark 1.1 ranked 3rd on the Debate Benchmark, trailing only Fable 5 and Opus 4.7, but outscoring GPT-5.6 Sol.

The cited post highlights performance gains across several models: GPT-5.6 Sol, Grok 4.5, Sonnet 5, and MiniMax-M3 all improved over their predecessors. However, Muse Spark 1.1 emerged as one of the most surprising high scorers with 1688 points.

Original post →

More from Models

Models channel →