Muse Spark 1.1 Leads in Benchmark Performance

alexandr_wang · x · 2026-07-13

Alexandr Wang shared new evaluation results from theoretical computer science/finite model theory: Muse Spark 1.1 outperforms Opus, Grok 4.5, and Gemini on this benchmark.

The cited benchmark specifies that models are given several small graphs and must output a first-order logic formula describing the properties of designated nodes across multiple graphs simultaneously. The evaluation consists of 64 questions, heavily focusing on inductive reasoning and formal expression capabilities.

Original post →

More from Models

Models channel →